Event date · · Diagram-MMU

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

FACT STATEMENT

Diagram-MMU is a multi-modal benchmark with 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on diagram-to-code parsing, diagram-to-code editing, and diagram question answering, plus agentic settings. Evaluation of 12 MLLMs shows diagram-to-code tasks are more challenging than diagram question answering; models reason well but struggle to parse and edit. Under agentic settings, most models improve parsing and editing but degrade on question answering, while Claude-4.6 Opus improves across all three tasks.

What happened

Researchers introduced Diagram-MMU, a benchmark for assessing multimodal large language models (MLLMs) on scientific diagram understanding. It contains 3.7k diagrams and 18.3k questions across six domains, testing diagram-to-code parsing, editing, and question answering. Testing 12 MLLMs revealed that diagram-to-code tasks are harder than question answering; models can reason over diagrams but struggle to parse and edit them. Agentic settings improved parsing and editing for most models but reduced question-answering performance, except Claude-4.6 Opus, which improved on all tasks.

Technical significance

The benchmark separates diagram reasoning from diagram-to-code generation, showing a capability gap: MLLMs can answer questions about diagrams but fail to produce accurate code (e.g., LaTeX TikZ). Agentic workflows help code generation but hurt question answering, suggesting a trade-off between tool use and direct reasoning. Claude-4.6 Opus is the only model that improves across all tasks, indicating stronger integration of agentic capabilities.

Industry impact

Scientific writing tools like OpenAI Prism rely on diagram-to-code conversion. This benchmark highlights that current MLLMs are not yet reliable for automatic diagram parsing and editing, limiting adoption in scientific collaboration software. The performance gap creates an opportunity for specialized models or fine-tuning to improve diagram-to-code generation.

Decision value

For companies building scientific writing or collaboration tools, this benchmark identifies a key technical bottleneck. Improving diagram-to-code accuracy could differentiate products like OpenAI Prism. The data and evaluation framework can guide R&D investment in multimodal code generation.

What to watch

Expect follow-up work on improving diagram-to-code generation, possibly through better training data, specialized architectures, or agentic pipelines. Watch for updates to Claude models and other MLLMs targeting scientific diagram tasks. The benchmark may become a standard evaluation for scientific AI assistants.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.