OpenBMB released MiniCPM5-2B-DSpark, a speculative decoding draft model for MiniCPM5-2B
OpenBMB released MiniCPM5-2B-DSpark on Hugging Face, a 5-layer draft model with 323,776,001 parameters trained to pair with MiniCPM5-2B. It is Apache-2.0 licensed and achieves an aggregate acceptance length of 5.5174 at temperature 0.
China context
- Original name
- 面壁智能
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can download the Apache-2.0 licensed weights from Hugging Face and integrate the draft model with SGLang for speculative decoding.
- For investors
- The release of a specialized draft model indicates OpenBMB's focus on inference efficiency for edge deployments, which may strengthen its position in the on-device AI market.
MiniCPM5-2B-DSpark is a speculative decoding draft checkpoint designed for exact pairing with MiniCPM5-2B and its tokenizer. It has 5 draft layers, 323,776,001 parameters, and predicts 7 tokens per forward pass using target hidden-state layers [1, 10, 20, 30, 39]. The model was trained on 7,054,154,509 tokens across 1,959,525 sequences for 6 epochs with a maximum sequence length of 12,288. Evaluation shows acceptance lengths of 6.0496 for math, 6.1106 for code, 4.1585 for general, and 5.5174 aggregate at temperature 0. The model is released under Apache-2.0 and can be used with SGLang.
The draft model uses a 5-layer architecture with 323,776,001 parameters and predicts 7 tokens per forward pass, targeting hidden-state layers [1, 10, 20, 30, 39] of MiniCPM5-2B. Training used a mixture of general-domain, mathematics, and code prompts with responses generated by MiniCPM5-2B, optimizing CE + L1 + confidence loss. Acceptance length is defined as total completion tokens divided by speculative verification steps, with natural EOS termination and max new tokens=4096.
Developers deploying MiniCPM5-2B can reduce inference latency by using this draft model for speculative decoding, lowering serving costs for on-device and resource-constrained applications.
The release provides a ready-to-use speculative decoding component that can improve throughput for MiniCPM5-2B deployments without additional training, under a permissive Apache-2.0 license.
Watch for integration of MiniCPM5-2B-DSpark into inference frameworks beyond SGLang and community benchmarks comparing speedups against other draft models.