提出新基准ScholScan,测试模型像人一样通读论文找矛盾
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
- 设计扫描式任务,要求模型通读全文找逻辑不一致
- 覆盖13个领域715篇论文,含1800个错误类型标注
- 发现现有检索增强方法对扫描任务无效,揭示模型短板
随着多模态大模型(MLLMs)快速发展,人工智能在文献检索和部分推理任务上已表现良好,成为研究人员的有力助手,但仍远未实现自主研究。当前学术论文推理工作主要局限于以预设目标为中心的搜索范式,依赖相关性检索进行推理,难以支持研究者式的全文档理解、推理与验证。为弥合这一差距,我们提出 extbf{ScholScan},一个面向学术论文扫描式推理的新基准。ScholScan引入扫描式任务设定,要求模型如人类研究者一般阅读并交叉核验整篇论文,扫描文档以识别一致性问题。该基准包含从9类错误中提取的1800个精心标注的问题,覆盖13个自然科学领域和715篇论文,并提供证据定位与推理路径的详细标注,以及统一评估协议。我们评估了15个模型在24种输入配置下的表现,并对所有错误类别进行了细粒度分析。结果显示,检索增强生成(RAG)方法并未带来显著提升,暴露了当前MLLM在扫描式任务上的系统性缺陷,凸显了ScholScan的挑战性。我们期待ScholScan成为扫描式任务范式的标杆之作。
原文摘要 · Abstract (English)
With the rapid progress of multimodal large language models (MLLMs), AI already performs well at literature retrieval and certain reasoning tasks, serving as a capable assistant to human researchers, yet it remains far from autonomous research. The fundamental reason is that current work on academic paper reasoning is largely confined to a search-oriented paradigm centered on pre-specified targets, with reasoning grounded in relevance retrieval, which struggles to support researcher-style full-document understanding, reasoning, and verification. To bridge this gap, we propose \textbf{ScholScan}, a new benchmark for academic paper reasoning. ScholScan introduces a scan-oriented task setting that asks models to read and cross-check entire papers like human researchers, scanning the document to identify consistency issues. The benchmark comprises 1,800 carefully annotated questions drawn from nine error categories across 13 natural-science domains and 715 papers, and provides detailed annotations for evidence localization and reasoning traces, together with a unified evaluation protocol. We assessed 15 models across 24 input configurations and conducted a fine-grained analysis of MLLM capabilities for all error categories. Across the board, retrieval-augmented generation (RAG) methods yield no significant improvements, revealing systematic deficiencies of current MLLMs on scan-oriented tasks and underscoring the challenge posed by ScholScan. We expect ScholScan to be the leading and representative work of the scan-oriented task paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。