Shanghai AI Laboratory open-sources ResOPD, a sparse on-policy distillation method
Shanghai AI Laboratory released ResOPD, an open-source repository implementing Tail Residualization for Sparse On-Policy Distillation, built on verl. The repository includes a processed DeepMath subset with 1,024 training samples and 252 held-out samples.
China context
- Original name
- 上海人工智能实验室
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can integrate ResOPD into verl-based distillation pipelines to reduce gradient variance without extra teacher forward passes, potentially lowering compute costs.
- For investors
- The open-source release of an efficient distillation method from a major Chinese lab may intensify competition in model compression, affecting the value proposition of proprietary distillation services.
Translated from Chinese. Quotes and facts link to the original sources.
ResOPD estimates the full-vocabulary reverse-KL gradient from the teacher's Top-k probabilities and the sampled-token score. It computes the observable coarse gradient exactly and samples only the unresolved within-tail residual, using the same sparse teacher payload and no additional teacher forward passes. The repository provides the core implementation and its integration with verl, along with a processed DeepMath subset for distillation training and held-out evaluation.
ResOPD separates an exact coarse update from a sampled fine-tail correction: it integrates the teacher Top-k contributions analytically, accounts for the tail bucket using residual student and teacher masses, and samples only the unresolved fine-tail correction on tail hits. The method is designed to reduce estimator variance compared to sampled-token on-policy distillation, with reported exact gradient audits showing 52.5–75.0% lower covariance trace than sampled-token OPD.
Developers using verl for on-policy distillation can adopt ResOPD to reduce gradient variance without additional teacher forward passes, potentially lowering compute cost for distilling large models into smaller ones. The open-source release under Apache 2.0 allows integration into existing verl pipelines.
ResOPD offers a method to improve the efficiency of on-policy distillation, which could reduce training costs for organizations distilling large language models into smaller, deployable models. The open-source Apache 2.0 license permits commercial use, potentially lowering barriers for enterprises seeking efficient model compression.
The next verifiable signal is whether the paper's reported benchmark results (50.12% average score on mathematical reasoning for Qwen3.5-27B → Qwen3.5-4B) are reproduced by independent evaluators using the released code and dataset. Adoption can be tracked via GitHub stars, forks, and issues on the InternLM/ResOPD repository.