构建多模态大模型科学论文综合推理评测基准
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

- 基于真实论文设计四类协同任务,评估多模态融合与推理能力
- 跨领域测试发现开源与闭源模型均存在系统性推理短板
- 适合研究科学智能、批判性思维评估的AI研究人员使用
理解科学论文不仅需要回答孤立问题或摘要内容,还需整合文本与视觉信息、解读实验证据、跨源信息融合,并批判性评估科学论断。现有评测通常孤立考察这些能力,难以衡量其作为整体认知功能的综合表现。本文提出PaperMind,一个面向多模态大模型的科学论文综合推理评测基准。数据源自农业、生物、化学、计算机科学、医学、物理和经济学七个领域的真实论文,包含四类互补任务:多模态定位、实验解读、跨源证据推理与批判性评估,全面刻画科学推理的认知维度。通过多任务分析,可诊断模型在集成推理中的行为模式。对开源与闭源多模态大模型的实测显示,各任务间存在持续性能差距,揭示了当前模型在综合科学推理与批判性评价方面的普遍挑战。评测数据集及代码已开源。
原文摘要 · Abstract (English)
Understanding scientific papers requires more than answering isolated questions or summarizing content. It involves an integrated reasoning process that grounds textual and visual information, interprets experimental evidence, synthesizes information across sources, and critically evaluates scientific claims. However, existing benchmarks typically assess these abilities in isolation, making it difficult to evaluate scientific paper understanding as a unified set of interacting cognitive abilities. In this work, we introduce PaperMind, a benchmark designed to evaluate integrated and agent-oriented scientific reasoning over research papers. PaperMind is constructed from real scientific papers across seven domains, including agriculture, biology, chemistry, computer science, medicine, physics, and economics. It comprises four complementary task families that collectively operationalize distinct cognitive facets of scientific paper reasoning, including multimodal grounding, experimental interpretation, cross-source evidence reasoning, and critical assessment. By analyzing model behavior across multiple tasks, PaperMind enables a diagnostic evaluation of integrated scientific reasoning behaviors that are difficult to assess through isolated task evaluations. Extensive experiments on both opensource and closed-source multimodal LLMs reveal consistent performance gaps across tasks, highlighting persistent challenges in integrated scientific reasoning and critique. Our benchmark and dataset are available at https:// github.com/Yanjun-Zhao/PaperMind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。