ModelBest open-sources Meshy, an asynchronous RL engine for LLMs
ModelBest (OpenBMB) released Meshy, an open-source asynchronous reinforcement learning engine for LLMs, on GitHub. It models every RL role as an independent service communicating through a TransferQueue data plane. The repository includes recipes for Qwen3 and MiniCPM5 models.
China context
- Original name
- 面壁智能
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can integrate Meshy into existing RL pipelines to experiment with asynchronous training without building custom infrastructure.
- For investors
- The open-sourcing of Meshy may reduce the market for commercial RL training platforms, affecting startups in that niche.
Meshy is an asynchronous RL engine where inference, training, and rollout run as independent services connected by a TransferQueue. It supports on-policy, bounded off-policy, and fully asynchronous training with the same services, differing only in rollout pacing. The framework is built on SGLang and torchtitan, and includes recipes for models such as Qwen3-1.7B, Qwen3-8B, Qwen3-30B-A3B, and MiniCPM5 variants.
Meshy eliminates a central driver by using queue columns as both data and control plane; column readiness is the only control signal. GPU ownership is passed as a token over TransferQueue, allowing arbitrary colocation of services. Topology is computed SPMD-style on each machine from a declarative recipe, with no service discovery.
Developers outside China can now use Meshy to run asynchronous RL training on their own GPU clusters without licensing fees, reducing the cost of building RL-tuned LLMs. The open-source release may pressure commercial RL infrastructure vendors to differentiate on managed services or support.
Meshy provides a free, open-source alternative to proprietary RL training infrastructure, potentially lowering the barrier for startups and researchers to experiment with asynchronous RL. Its colocation and topology features may reduce GPU idle time and simplify multi-node orchestration.
A specific signal to check is whether Meshy gains adoption in open-source RL training pipelines, such as forks or integrations with popular frameworks like TRL or OpenRLHF. Another observable is the release of trained checkpoints or benchmark results from the Meshy recipes.