Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
A paper published on arXiv on 2026-08-31 studies how large reasoning models can improve as human supervision recedes, examining reward and experience axes through a five-level ladder from L0 to L4.
The paper explores scaling large reasoning models beyond human supervision by analyzing two dimensions: the reward axis (from per-instance human judgments to reusable verifiers and autonomous rewards) and the experience axis (from human-curated tasks to self-generated curricula and co-evolution). It introduces a five-level ladder (L0-L4) to identify which parts of the learning process remain under human control and highlights risks of increasingly autonomous rewards.
The paper proposes a structured framework for reducing human supervision in reinforcement learning with verifiable rewards (RLVR), extending beyond math and code to open-ended and agentic tasks. Key technical signals include the development of reusable verifiers, self-generated curricula, and autonomous co-evolution of environments and models.
This research signals a shift toward more autonomous training pipelines, potentially reducing the need for human feedback at scale. It may influence how AI labs design training infrastructure and could accelerate progress toward more general reasoning capabilities.
Reducing human supervision in training could lower costs and speed up model iteration, creating competitive advantages for organizations that can safely implement autonomous reward and experience generation.
Observable next signals include follow-up papers implementing the L0-L4 ladder, open-source tooling for autonomous reward design, and industry adoption of self-generated curricula in large model training.