Shanghai AI Laboratory releases SciDocBench benchmark for scientific document understanding
Shanghai AI Laboratory released SciDocBench, a workflow-centered benchmark for scientific document understanding, on 2026-09-04. It includes 124 manually designed questions, 496 evaluation instances, and is integrated into VLMEvalKit.
China context
- Original name
- 上海人工智能实验室
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can use SciDocBench via VLMEvalKit to evaluate their models on scientific document understanding tasks, with the benchmark data publicly available on Hugging Face.
- For investors
- The release of SciDocBench by Shanghai AI Laboratory signals continued investment in evaluation infrastructure for scientific AI, which may influence funding decisions for startups in the document AI space.
SciDocBench is a benchmark for scientific document understanding that evaluates evidence-grounded operations such as locating support, verifying numerical relations, and cross-paper synthesis. It contains 124 manually designed questions across 7 capability groups and 19 subtasks, covering 5 scientific domains in English and Chinese. The benchmark is available on Hugging Face and integrated into VLMEvalKit for standardized evaluation.
The benchmark uses 496 matched evaluation instances across four settings (EN-AF, EN-IL, ZH-AF, ZH-IL) and three evaluator families: rule-based, LLM-as-a-judge, and execution-based. The leaderboard shows Claude Opus 5 leading with an overall score of 62.60, followed by GPT-5.6-Sol (61.00) and Gemini 3.6 Flash (59.87). No evaluated model reaches 63, indicating room for improvement in scientific document understanding.
Researchers and developers working on scientific document assistants now have a standardized benchmark to measure and compare model performance on evidence-grounded tasks, which may shift development priorities toward improving cross-document reasoning and verification capabilities.
For companies building scientific document processing tools, SciDocBench provides a concrete evaluation framework to benchmark their models against leading commercial and open-source models, potentially informing product development and positioning.
The next observable signal is whether model providers release updated versions targeting higher SciDocBench scores, particularly in the ZH-AF and ZH-IL settings where current top models score below 66. The benchmark's integration into VLMEvalKit may lead to broader adoption in academic and industry evaluations.