Event date · · Shanghai AI Laboratory

Shanghai AI Laboratory open-sources ResOPD, a sparse on-policy distillation method

FACT STATEMENT

Shanghai AI Laboratory released ResOPD, an open-source repository implementing Tail Residualization for Sparse On-Policy Distillation, built on verl. The repository includes a processed DeepMath subset with 1,024 training samples and 252 held-out samples.

China context

Original name
上海人工智能实验室
Outside China
Open weights · github.com
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can integrate ResOPD into verl-based distillation pipelines to reduce gradient variance without extra teacher forward passes, potentially lowering compute costs.
For investors
The open-source release of an efficient distillation method from a major Chinese lab may intensify competition in model compression, affecting the value proposition of proprietary distillation services.

Translated from Chinese. Quotes and facts link to the original sources.

What happened

ResOPD estimates the full-vocabulary reverse-KL gradient from the teacher's Top-k probabilities and the sampled-token score. It computes the observable coarse gradient exactly and samples only the unresolved within-tail residual, using the same sparse teacher payload and no additional teacher forward passes. The repository provides the core implementation and its integration with verl, along with a processed DeepMath subset for distillation training and held-out evaluation.

Technical significance

ResOPD separates an exact coarse update from a sampled fine-tail correction: it integrates the teacher Top-k contributions analytically, accounts for the tail bucket using residual student and teacher masses, and samples only the unresolved fine-tail correction on tail hits. The method is designed to reduce estimator variance compared to sampled-token on-policy distillation, with reported exact gradient audits showing 52.5–75.0% lower covariance trace than sampled-token OPD.

Industry impact

Developers using verl for on-policy distillation can adopt ResOPD to reduce gradient variance without additional teacher forward passes, potentially lowering compute cost for distilling large models into smaller ones. The open-source release under Apache 2.0 allows integration into existing verl pipelines.

Decision value

ResOPD offers a method to improve the efficiency of on-policy distillation, which could reduce training costs for organizations distilling large language models into smaller, deployable models. The open-source Apache 2.0 license permits commercial use, potentially lowering barriers for enterprises seeking efficient model compression.

What to watch

The next verifiable signal is whether the paper's reported benchmark results (50.12% average score on mathematical reasoning for Qwen3.5-27B → Qwen3.5-4B) are reproduced by independent evaluators using the released code and dataset. Adoption can be tracked via GitHub stars, forks, and issues on the InternLM/ResOPD repository.

Latest in Chinese AI

  1. XiaomiReportedly: Xiaomi retires MiMo-V2.5-Pro and MiMo-V2.5, auto-switching to V2.6 models on October 14
  2. ByteDanceReportedly: Doubao Work adds canvas feature and updates models
  3. HuaweiReportedly: Huawei launches Smart Door Lock 2 Pro Enjoy Edition with AI 3D face recognition 3.0
  4. Reportedly: Schneider Electric acquires PTC for $22.6B
  5. ByteDanceReportedly: Doubao Work adds an infinite creation canvas with Seedream 5.0 Flash and Doubao 2.1 Lite

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.