Xiaomi MiMo released MiMo-V2.6-Flash-MOPD on Hugging Face
Xiaomi MiMo released MiMo-V2.6-Flash-MOPD on Hugging Face, a sparse Mixture-of-Experts model with 309B total and 15B activated parameters, 1M-token context, and multimodal support for text, image, video, and audio. The model is available under an MIT license and can be deployed via SGLang or vLLM.
China context
- Original name
- 小米 MiMo
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can download the model weights from Hugging Face and deploy locally using SGLang or vLLM, with an MIT license permitting commercial use and modification.
- For investors
- Xiaomi's release of a 309B-parameter open-weight model signals continued investment in frontier AI and a strategy to build an ecosystem around MiMo, which can be tracked through subsequent model releases and developer adoption metrics.
Xiaomi MiMo released MiMo-V2.6-Flash-MOPD on Hugging Face. The model is a sparse Mixture-of-Experts (MoE) architecture with 309B total parameters and 15B activated parameters, supporting a 1M-token context length and multimodal inputs including text, image, video, and audio. It is an upgrade of the MiMo-V2.6-Flash-RL checkpoint, incorporating MOPD2 (Multi-Objective Policy Distillation) to fuse domain-specialized teachers and mitigate tool-call repetition in agentic settings. The model is released under the MIT license and can be deployed using SGLang or vLLM, with recommended sampling parameters of temperature=1.0 and top_p=0.95. It is also available through Xiaomi MiMo's API platform, AI Studio, MiMo Code, Xiaomi MiMo Desktop, and OpenRouter.
MiMo-V2.6-Flash-MOPD uses a sparse MoE architecture with 256 routed experts, activating 8 per token, and a hybrid attention pattern combining sliding window attention (SWA) and global attention (GA). The vision encoder is a 681M-parameter MiMo ViT with 28 layers (24 SWA + 4 GA), and the audio encoder comprises a 308M AudioTokenizer and a 127M audio patch encoder. A 5-layer speculative decoder (MTP) predicts 7 subsequent tokens per forward pass for parallel verification. The MOPD2 training stage distills multiple domain-specialized teachers on-policy, including mixRL teachers for verifiable tasks and SFT teachers for open-domain tasks, to improve performance in long-horizon game development, scientific research, and embodied intelligence, and to reduce tool-call repetition.
Developers outside China can now access a 309B-parameter multimodal MoE model under an MIT license, enabling local deployment and fine-tuning without API costs or data-sharing constraints. This release intensifies open-weights competition by providing a high-capacity model with a 1M-token context and agentic tool-calling improvements, directly challenging proprietary API offerings.
For enterprises, the MIT license permits commercial use and modification, reducing procurement risk and enabling on-premise deployment for data-sensitive applications. The 1M-token context and multimodal capabilities address use cases in document analysis, video understanding, and agentic workflows, while the sparse MoE design offers a balance between capability and inference cost.
Observable next signals include adoption metrics on Hugging Face (downloads, likes, community forks), integration into third-party inference frameworks beyond SGLang and vLLM, and independent benchmarks evaluating the model's tool-call repetition rate and long-horizon task performance. Xiaomi's continued release of MOPD checkpoints for both Pro and Flash variants suggests a pattern of iterative open-weight updates.