评测大模型在文档中跨图表与文本的多跳推理能力
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

- 基于逻辑生成框架构建可控复杂度的图文推理题
- 模型最高仅62.83%准确率,远低于人类超90%表现
- 适合评估多跳推理、跨模态理解能力的模型
多模态大语言模型在结构化视觉理解任务(如图表和文档问答)中表现强劲。然而现有基准通常孤立评估这些领域,未充分探索关键能力:模型能否利用文本上下文判断图表证据的选择、解读与整合方式。我们提出DocHop,一个面向文档类图像中图表-文本联合推理的基准。在DocHop中,文档叙述设定多步组合约束,图表提供对应数据值;问题基于叙述中的语义参考标签,要求模型先从上下文定位目标实体,再跨多张图表聚合证据。为实现系统性评估,我们通过随机逻辑优先生成管道构建了涵盖6类任务的2,074个样例。对多种专有及开源MLLM的实验表明,模型与人类表现存在显著差距:标注者准确率超90%,最佳模型仅达62.83%。增强推理能力的模型表现更优,但性能随推理复杂度上升而下降。总体而言,DocHop为挑战性的多跳文档推理提供了可控测试平台。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。