arXiv:2603.07931cs.CL2026-03被引 3

评测大模型在长科学文档中多跳推理能力,强调证据整合与定位。

BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence

  • 构建支持链式与发散结构的多跳推理数据集
  • 发现现有模型在证据聚合与定位上存在系统性缺陷
  • 适合研究模型推理机制与可解释性的研究人员

多跳问答广泛用于评估大语言模型的推理能力,但现有基准大多仅关注最终答案正确性,忽视中间推理过程,尤其在长篇多模态文档中更为明显。我们提出BRIDGE,一个面向长科学论文的多跳推理基准,要求整合文本、表格和图表中的证据。该数据集支持链式与发散结构,并提供逐步推理标注,实现对推理过程的细粒度评估。基于顶尖LLM与多模态检索增强生成系统实验表明,现有模型在证据聚合与定位方面存在系统性不足,这些缺陷在仅评价答案正确性的传统评估中难以察觉。BRIDGE为诊断长多模态文档中的推理失败提供了精准测试平台。

原文摘要 · Abstract (English)

Multi-hop question answering (QA) is widely used to evaluate the reasoning capabilities of large language models, yet most benchmarks focus on final answer correctness and overlook intermediate reasoning, especially in long multimodal documents. We introduce BRIDGE, a benchmark for multi-hop reasoning over long scientific papers that require integrating evidence across text, tables, and figures. The dataset supports both chain-like and fan-out structures and provides explicit multi-hop reasoning annotations for step-level evaluation beyond answer accuracy. Experiments with state-of-the-art LLMs and multimodal retrieval-augmented generation (RAG) systems reveal systematic deficiencies in evidence aggregation and grounding that remain hidden under conventional answer-only evaluation. BRIDGE provides a targeted testbed for diagnosing reasoning failures in long multimodal documents.

多跳推理模型评估科学文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。