构建科学图表多模态评测集,评估大模型解析与编辑图表能力。
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

- 设计三类任务:图表转代码、图表编辑、图表问答,支持智能体设置。
- 涵盖3.7千张图表、1.8万道人工验证问题,覆盖六个科学领域。
- 发现模型理解图表能力强但生成代码差,仅Claude-4.6 Opus全任务提升。
多模态大语言模型(MLLMs)在科学写作与协作中展现出日益增强的能力,如OpenAI Prism提供免费科学写作协作空间,并支持将科学图表直接转换为LaTeX TikZ代码。本文构建了Diagram-MMU——一个面向科学图表解析与理解的多模态基准评测集。该基准包含3.7k张精心筛选的图表和18.3k条人工验证的问题,覆盖六个科学领域。评估聚焦三类常见于科研工作流的任务:图表转代码、图表转代码编辑、图表问答,并在每项任务中引入代理式(agentic)设置。对12个MLLMs的评估显示,图表转代码任务比图表问答更具挑战性:模型虽能较好推理图表内容,但在解析与编辑方面表现不佳,凸显了提升其生成代码能力的必要性。在代理设置下,多数模型在解析与编辑任务中表现改善,但在问答任务中性能下降;而Claude-4.6 Opus则在所有三项任务中均实现持续提升。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks. Project Page: https://vi-ocean.github.io/projects/diagram-mmu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。