Xiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face
Model: MiMo-V2.6-Pro · availability, license and releases
Xiaomi MiMo released MiMo-V2.6-Pro-RL, a 1.02T-parameter sparse MoE model with 42B activated parameters, on Hugging Face under an MIT license. It supports 1M-token context and text, image, video, and audio modalities.
China context
- Original name
- 小米 MiMo
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders can download and fine-tune the model under MIT license, with 1M-token context and multimodal support for agentic applications.
- For investors
- Xiaomi's release of a frontier-scale open-weights model signals its commitment to competing in the global AI model market, which may pressure other open-weights providers.
MiMo-V2.6-Pro-RL is the flagship checkpoint of the MiMo-V2.6 series, built to scale reinforcement learning toward self-improvement. It uses a sparse Mixture-of-Experts architecture with 1.02T total parameters and 42B activated parameters, a 1M-token context length, and native omnimodal support for text, image, video, and audio. The model is available on Hugging Face and ModelScope under an MIT license.
The model employs a sparse MoE architecture with 384 routed experts (8 activated) and a 681M-parameter vision encoder. It uses Group Relative Policy Optimization (GRPO) with 1,568 prompts × 16 rollouts per step and a groupwise agentic grading system (GRS and GAR) to scale reinforcement learning. A 5-layer multi-token prediction speculative decoder is included.
Developers outside China can now access and modify a frontier-scale open-weights model with 1M-token context and omnimodal capabilities under a permissive MIT license, reducing the cost and complexity of building advanced AI applications without relying on proprietary APIs.
The MIT license allows commercial use and modification, enabling enterprises to deploy a high-capacity multimodal model without licensing fees. The 1M-token context and agentic capabilities target long-horizon tasks such as repository-level coding and multi-session agents.
Watch for whether MiMo-V2.6-Pro-RL's benchmark results are independently reproduced and whether the model gains adoption in agentic coding and cybersecurity tasks, where it shows competitive scores against Claude Opus 5 and GPT-5.6.