首个评估模型理解科学图示能力的基准,揭示现有模型与人类差距
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
- 构建1500个专家标注的科学图示问答数据集
- 18个主流多模态模型在该任务上表现远低于人类
- 适合关注科学文献理解的AI研究者使用
本文提出MISS-QA,首个专门用于评估模型理解科学文献中示意图能力的基准。该基准包含465篇科学论文中的1500个专家标注的问答对。模型需基于论文整体上下文,解读展示研究概览的示意图并回答信息查询问题。我们评估了18个前沿多模态基础模型(包括o4-mini、Gemini-2.5-Flash、Qwen2.5-VL)的表现,发现这些模型在MISS-QA上的表现显著落后于人类专家。通过对不可回答问题的分析及详细的错误归因,进一步揭示了当前模型在理解多模态科学文献时的优势与局限,为提升模型能力提供了关键洞见。
原文摘要 · Abstract (English)
This paper introduces MISS-QA, the first benchmark specifically designed to evaluate the ability of models to interpret schematic diagrams within scientific literature. MISS-QA comprises 1,500 expert-annotated examples over 465 scientific papers. In this benchmark, models are tasked with interpreting schematic diagrams that illustrate research overviews and answering corresponding information-seeking questions based on the broader context of the paper. We assess the performance of 18 frontier multimodal foundation models, including o4-mini, Gemini-2.5-Flash, and Qwen2.5-VL. We reveal a significant performance gap between these models and human experts on MISS-QA. Our analysis of model performance on unanswerable questions and our detailed error analysis further highlight the strengths and limitations of current models, offering key insights to enhance models in comprehending multimodal scientific literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。