Shanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification
Shanghai AI Laboratory released AdvancedMathBench, a benchmark suite with 245 proof-generation problems and 888 annotated proof trajectories, plus a trained AutoVerifier. The repository is available on GitHub and includes evaluation results for models like GPT-5.5-xhigh and DeepSeek-V4-Pro.
China context
- Original name
- 上海人工智能实验室
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers can use the benchmark and AutoVerifier to evaluate their models' mathematical proof capabilities, with the repository providing code and data for reproduction.
- For investors
- The release indicates Shanghai AI Laboratory's continued investment in AI for mathematics, which may be relevant for assessing the lab's research direction and potential commercial applications.
AdvancedMathBench evaluates language models on constructing and verifying natural-language proofs in advanced mathematics. It includes ProverBench (245 expert-reviewed problems) and VerifierBench (888 proof trajectories with expert annotations). The default pessimistic protocol accepts a proof only when all eight judgments accept it. Main results show GPT-5.5-xhigh achieving 64.5% on ProverBench UG and 48.9% on QE, while DeepSeek-V4-Pro achieves 54.0% and 40.0% respectively.
The benchmark separates proof generation from proof verification, using a pessimistic acceptance protocol and meta-verification to assess error analysis agreement. AutoVerifier is a trained proof verifier for automated evaluation. The repository requires Python 3.10+ and provides an API evaluator with no internal dependencies.
AI model developers and researchers evaluating mathematical reasoning now have a standardized benchmark that distinguishes proof generation from verification, which can change how models are compared and improved for advanced mathematics.
For organizations developing or using AI for mathematical reasoning, this benchmark provides a tool to assess and compare model capabilities in proof generation and verification, which could inform model selection and development priorities.
Observable next signals include adoption of AdvancedMathBench in model evaluations, updates to the benchmark, and improvements in model performance on proof verification tasks.