Event date · · Shanghai AI Laboratory

Shanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification

FACT STATEMENT

Shanghai AI Laboratory released AdvancedMathBench, a benchmark suite with 245 proof-generation problems and 888 annotated proof trajectories, plus a trained AutoVerifier. The repository is available on GitHub and includes evaluation results for models like GPT-5.5-xhigh and DeepSeek-V4-Pro.

China context

Original name
上海人工智能实验室
Outside China
Open weights · github.com
Claims
Company-reported; not yet independently evaluated
For builders
Developers can use the benchmark and AutoVerifier to evaluate their models' mathematical proof capabilities, with the repository providing code and data for reproduction.
For investors
The release indicates Shanghai AI Laboratory's continued investment in AI for mathematics, which may be relevant for assessing the lab's research direction and potential commercial applications.
What happened

AdvancedMathBench evaluates language models on constructing and verifying natural-language proofs in advanced mathematics. It includes ProverBench (245 expert-reviewed problems) and VerifierBench (888 proof trajectories with expert annotations). The default pessimistic protocol accepts a proof only when all eight judgments accept it. Main results show GPT-5.5-xhigh achieving 64.5% on ProverBench UG and 48.9% on QE, while DeepSeek-V4-Pro achieves 54.0% and 40.0% respectively.

Technical significance

The benchmark separates proof generation from proof verification, using a pessimistic acceptance protocol and meta-verification to assess error analysis agreement. AutoVerifier is a trained proof verifier for automated evaluation. The repository requires Python 3.10+ and provides an API evaluator with no internal dependencies.

Industry impact

AI model developers and researchers evaluating mathematical reasoning now have a standardized benchmark that distinguishes proof generation from verification, which can change how models are compared and improved for advanced mathematics.

Decision value

For organizations developing or using AI for mathematical reasoning, this benchmark provides a tool to assess and compare model capabilities in proof generation and verification, which could inform model selection and development priorities.

What to watch

Observable next signals include adoption of AdvancedMathBench in model evaluations, updates to the benchmark, and improvements in model performance on proof verification tasks.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  3. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  4. Manus AIManus AI launches Manus Flex, letting users bring their own API key into the Manus workspace
  5. AlibabaAlibaba DAMO Academy Unveils DAMO EAGLE AI Model for Early Esophageal Cancer Detection

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.