Event date · · Google Antigravity

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

FACT STATEMENT

Stellar Colosseum is a model-agnostic harness for allocating inference across research in mathematics and theoretical computer science. It explores alternative strategies before proof construction, uses a readiness gate to decide when a route is mature enough to decompose, represents the proof plan as interdependent section-level subproblems, and routes verifier findings back to the affected part of the argument. It generates candidates in parallel, attacks them with targeted falsification, and combines candidates and their critiques into a single research artifact through overlapping random-sample tree aggregation. The Colosseum workflow has been integrated into Google Antigravity's Teamwork framework as the Long Proof pattern. Using Colosseum with Gemini 3.1 Pro, the authors obtain several new results.

What happened

Stellar Colosseum is a many-agent harness designed for long-horizon research in mathematics and theoretical computer science. It addresses the unreliability of language models on complex problems by allocating inference across multiple stages: exploring alternative strategies, using a readiness gate before decomposition, representing proof plans as interdependent subproblems, and routing verifier feedback to relevant sections. The system generates candidates in parallel, applies targeted falsification, and aggregates results via overlapping random-sample tree aggregation. The workflow is integrated into Google Antigravity's Teamwork framework as the Long Proof pattern. Evaluations on theorem-proving and competitive programming benchmarks, using Gemini 3.1 Pro, demonstrate new results.

Technical significance

The harness introduces a readiness gate to control decomposition timing, section-level subproblem representation for proof plans, and overlapping random-sample tree aggregation to combine candidates and critiques. These mechanisms aim to improve reliability on long-horizon tasks by managing uncertainty and interdependence across decisions.

Industry impact

Integration into Google Antigravity's Teamwork framework suggests a path toward productizing multi-agent research workflows for mathematical and theoretical computer science domains, potentially influencing enterprise and developer tools for complex reasoning.

Decision value

The harness could reduce the cost and time of long-horizon research by improving the reliability of language model outputs in mathematics and theoretical computer science, with potential applications in automated theorem proving and competitive programming.

What to watch

Observable next signals include publication of the new results obtained with Gemini 3.1 Pro, further benchmark evaluations, and potential adoption or extension of the Long Proof pattern in other Google Antigravity Teamwork applications.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.