构建科学实验图像评测基准,检验AI理解复杂科研图示的能力
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

- 从1084张专家精选图像中生成4264个问答对,覆盖多维度视觉感知
- 平均每幅图含14.3个面板,测试模型跨面板关系理解能力
- 评估模型在五类实验范式中的推理水平,发现当前AI远未达专家水准
我们提出SPUR,一个涵盖科学实验图像感知、理解与推理的综合性评测基准,包含4,264个基于1,084张专家筛选图像的问答对。该基准有三大创新:(1) 面板级细粒度感知,从数值、形态和信息定位三个维度评估多模态大语言模型(MLLMs)对六类细粒度面板的视觉感知能力;(2) 跨面板关系理解,利用平均每个样本含14.3个面板的复杂图像,测试MLLMs解析复杂跨面板关联的能力;(3) 专家级推理,评估模型在五种实验范式下的定性与定量推理能力,判断其是否能像人类专家一样从证据推导结论。对20个MLLMs及四种多模态思维链(MCoT)方法的全面评估显示,当前模型在科学图像解读方面仍远未达到专家水平,凸显人工智能在科学领域(AI4S)研究中的关键瓶颈。
原文摘要 · Abstract (English)
We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images. SPUR features three key innovations: (1) Panel-Level Fine-Grained Perception: evaluating the visual perception of multimodal large language models (MLLMs) across three dimensions (numerical, morphological, and information localization) on six fine-grained panel types; (2) Cross-Panel Relation Understanding: utilizing complex images with an average of 14.3 panels per sample to evaluate MLLMs' ability to decipher intricate cross-panel relations; (3) Expert-Level Reasoning: assessment of qualitative and quantitative reasoning across five experimental paradigms to determine if models can infer conclusions from evidence as human experts do. Comprehensive evaluation of 20 MLLMs and four multimodal Chain-of-Thought (MCoT) methods reveals that current models fall significantly short of the expert-level requirements for scientific image interpretation, underscoring a critical bottleneck in AI for Science (AI4S) research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。