Xiaomi MiMo released MiMo-V2.6-Flash-RL on Hugging Face
Xiaomi MiMo released MiMo-V2.6-Flash-RL, a sparse Mixture-of-Experts model with 309B total and 15B activated parameters, on Hugging Face under the MIT license. It supports 1M-token context and text, image, video, and audio modalities.
China context
- Original name
- 小米 MiMo
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can download the weights from Hugging Face and integrate the model into agentic applications with 1M context and multimodal inputs.
- For investors
- Xiaomi's release of a 309B open-weights model under MIT license intensifies competition in the open-weights segment, which may pressure pricing and differentiation for other labs.
MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series, built to scale reinforcement learning toward self-improvement. It uses a sparse MoE architecture with 309B total parameters and 15B activated, a 1M-token context length, and native omnimodal support for text, image, video, and audio. The model was trained with a single mixed RL run across coding, general agents, visual, and cybersecurity tasks, using asynchronous Group Relative Policy Optimization on large batches. It also employs groupwise agentic grading for self-improvement and multi-prefix multi-teacher on-policy distillation. The model is available on Hugging Face and ModelScope under the MIT license.
The model uses a sparse MoE with 256 routed experts, activating 8 per token, and a 48-layer backbone with 39 sliding-window attention layers and 9 global attention layers. It includes a 681M-parameter vision encoder and a 308M audio tokenizer plus 127M audio patch encoder. Training used 1,568 prompts × 16 rollouts per step with billions of tokens per update. The RL process mixes tasks and harnesses in the same batch, and uses groupwise reward synthesis and advantage redistribution to rank passing trajectories.
Developers outside China can now use a 309B-parameter open-weights model with 1M context and omnimodal input under the MIT license, reducing the cost of building long-horizon agents compared to proprietary alternatives.
The MIT license allows commercial use, modification, and redistribution, enabling enterprises to deploy the model without licensing fees. The 15B activated parameters may offer a balance between capability and inference cost for agentic workloads.
Watch for whether the model's benchmark results are independently reproduced and whether the RL training recipe is adopted by other open-weights labs. The next observable signal is community fine-tunes or deployments on Hugging Face.