Event date · · arXiv

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

FACT STATEMENT

A paper on arXiv (cs.AI) presents an empirical study of data efficiency and data selection in On-Policy Distillation (OPD). The study finds that 1-shot OPD is consistently effective across all sampled training examples, with harder examples often yielding superior performance gain. The improvement is driven by longer chain-of-thought (CoT) paths rather than high token entropy, and training on longer CoT helps maintain closer alignment with the teacher and learn critical thinking patterns like reflection. The paper proposes a data selection method that selects only hard examples for training, including 'unsolvable' examples.

What happened

The paper investigates data-centric mechanisms in On-Policy Distillation (OPD), a post-training paradigm for enhancing large language models in reasoning domains. It reports that even training on a single example (1-shot OPD) is effective, and harder examples tend to produce better gains. The key driver is longer chain-of-thought paths, which enable the student to align with the teacher over longer reasoning horizons and learn patterns such as reflection. Based on these insights, the authors propose selecting only hard examples for training.

Technical significance

The finding that longer CoT paths, rather than token entropy, drive student improvement suggests that data selection should prioritize examples that elicit extended reasoning. The proposed method of selecting hard examples, including unsolvable ones, may improve data efficiency in OPD. Future work could explore scaling this approach and its interaction with different teacher-student architectures.

Industry impact

This research could influence how AI labs curate training data for post-training distillation, potentially reducing data requirements and costs while maintaining or improving reasoning performance. It may lead to more efficient fine-tuning pipelines for reasoning-focused models.

Decision value

Improved data efficiency in OPD can lower training costs and accelerate model development cycles for companies building reasoning-capable LLMs. The method may also enable smaller teams to achieve competitive performance with limited data.

What to watch

If the proposed data selection method proves robust across models and tasks, it could become a standard practice in OPD. Further research may examine the limits of using unsolvable examples and the generalizability to other domains beyond reasoning.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.