arXiv:2604.11307cs.AI2026-04ACL被引 1

构建首个多模态多文档科学推理评测基准,评估大模型跨论文深度研究能力。

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers

  • 基于2000+篇论文知识图谱,支持多文档科学问答与推理。
  • 通过语义密度采样确保任务复杂性,覆盖2000+个问答对。
  • 适合评估科研代理系统在长上下文、多源信息融合中的表现。

利用多模态大语言模型(MLLMs)加速前沿科学研究前景广阔,但如何严格评估此类系统仍不明确。现有评测基准主要聚焦单文档理解,而真实科研流程需整合多篇论文的文本、表格和图表等多模态证据。因此,多模态、多文档科学推理仍处于探索阶段且缺乏系统评估。为此,我们提出PaperScope,一个面向智能体深度科研的多模态多文档评测基准。PaperScope具备三大优势:(1) 结构化科学基础:基于涵盖三年内2000余篇人工智能论文的知识图谱,为研究型问题提供结构化支撑;(2) 语义密集证据构建:整合语义相关的关键信息节点,并采用优化的随机游走文章选择器,生成主题连贯的论文集合,保障足够的语义密度与任务复杂度;(3) 多任务科学推理评估:包含超过2000个问答对,覆盖推理、检索、摘要和问题求解,支持多步科学推理的综合评估。实验表明,即使先进系统如OpenAI Deep Research和Tongyi Deep Research在PaperScope上得分有限,凸显长上下文检索与深层多源推理的挑战。PaperScope不仅提供严谨评测基准,还配套可扩展的流水线,用于构建大规模多模态、多源深度研究数据集。

原文摘要 · Abstract (English)

Leveraging Multi-modal Large Language Models (MLLMs) to accelerate frontier scientific research is promising, yet how to rigorously evaluate such systems remains unclear. Existing benchmarks mainly focus on single-document understanding, whereas real scientific workflows require integrating evidence from multiple papers, including their text, tables, and figures. As a result, multi-modal, multi-document scientific reasoning remains underexplored and lacks systematic evaluation. To address this gap, we introduce PaperScope, a multi-modal multi-document benchmark designed for agentic deep research. PaperScope presents three advantages: (1) Structured scientific grounding. It is built on a knowledge graph of over 2,000 AI papers spanning three years, providing a structured foundation for research-oriented queries. (2) Semantically dense evidence construction. It integrates semantically related key information nodes and employs optimized random-walk article selector to sample thematically coherent paper sets, thereby ensuring adequate semantic density and task complexity. (3) Multi-task evaluation of scientific reasoning. It contains over 2,000 QA pairs across reasoning, retrieval, summarization, and problem solving, enabling evaluation of multi-step scientific reasoning. Experimental results show that even advanced systems such as OpenAI Deep Research and Tongyi Deep Research achieve limited scores on PaperScope, highlighting the difficulty of long-context retrieval and deep multi-source reasoning. PaperScope thus provides a rigorous benchmark alongside a scalable pipeline for constructing large-scale multi-modal, multi-source deep research datasets.

多模态科学推理评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。