Event date · · arXiv

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

FACT STATEMENT

A paper on arXiv (2609.01573v1) frames SFT-RL annotation budget allocation in terms of near-optimality, showing the near-optimal region is wide, widens with model scale, and transfers from small proxy models to large target models. Results hold across tasks, model families, and both preference-based off-policy and reward-supervision on-policy RL methods.

What happened

The paper addresses how to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training. It introduces a near-optimality framework, characterizing the set of allocations within a specified tolerance of peak performance. Empirically, the near-optimal region is wide even for small tolerances (2-10%), widens with model scale, and transfers reliably from small proxy models to large target models. This suggests small proxy-model experiments can identify a transferable near-optimal region, avoiding exhaustive large-scale search. The findings are consistent across tasks, model families, and both preference-based off-policy and reward-supervision on-policy RL methods. The paper also analyzes how asymmetry in annotation costs between SFT and RL data shifts the near-optimal region.

Technical significance

The near-optimal region widens with model scale, implying larger models are more robust to suboptimal SFT-RL allocation. Transferability from small proxy models to large target models enables efficient budget allocation without large-scale experimentation. The analysis of annotation cost asymmetry suggests that cost differences can shift the near-optimal region, potentially favoring one method over the other depending on relative costs.

Industry impact

This research provides a practical strategy for LLM developers to optimize post-training annotation budgets using small-scale experiments, reducing computational and financial costs. It may influence how AI labs allocate resources between SFT and RL, especially as model sizes grow and annotation costs vary.

Decision value

The ability to determine near-optimal SFT-RL allocation using small proxy models can significantly reduce the cost and time required for LLM post-training, enabling faster iteration and more efficient use of annotation budgets. This is particularly valuable for organizations developing large language models with limited resources.

What to watch

Future work may explore whether the near-optimal region transfer holds across different model architectures and more diverse RL algorithms. The framework could be extended to multi-stage post-training pipelines or combined with other budget allocation decisions, such as pretraining vs. post-training data.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.