构建多模态多文档科学问答数据集,评估大模型跨文献理解能力
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
- 基于70个论文簇构建多文档多模态问答数据集
- 18个大模型在跨文档推理上显著落后于人类专家
- 适合关注科学文献智能分析的研究者使用
现有大模型评估基准主要聚焦单文档、纯文本任务,难以全面反映科研工作流的复杂性——后者通常涉及非文本数据解读及跨多篇文献的信息整合。为此,我们提出M3SciQA,一个面向大模型综合评估的多模态、多文档科学问答基准。该数据集包含1,452个专家标注的问题,覆盖70个自然语言处理论文簇,每个簇包含主论文及其所有引用文献,模拟阅读单篇论文时需融合多源异构信息的真实流程。我们对18个大模型进行了全面评估,结果表明当前模型在多模态信息检索与跨文档推理方面仍显著弱于人类专家。此外,我们探讨了这些发现对大模型应用于多模态科学文献分析的启示。
原文摘要 · Abstract (English)
Existing benchmarks for evaluating foundation models mainly focus on single-document, text-only tasks. However, they often fail to fully capture the complexity of research workflows, which typically involve interpreting non-textual data and gathering information across multiple documents. To address this gap, we introduce M3SciQA, a multi-modal, multi-document scientific question answering benchmark designed for a more comprehensive evaluation of foundation models. M3SciQA consists of 1,452 expert-annotated questions spanning 70 natural language processing paper clusters, where each cluster represents a primary paper along with all its cited documents, mirroring the workflow of comprehending a single paper by requiring multi-modal and multi-document data. With M3SciQA, we conduct a comprehensive evaluation of 18 foundation models. Our results indicate that current foundation models still significantly underperform compared to human experts in multi-modal information retrieval and in reasoning across multiple scientific documents. Additionally, we explore the implications of these findings for the future advancement of applying foundation models in multi-modal scientific literature analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。