arXiv:2506.15594cs.CLcs.AI2025-06ACL被引 11

评测模型看懂表格图表的跨模态推理能力,发现大模型在长文档中表现大幅下降。

WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts

  • 构建1000道多选题,测试模型融合表格图表信息进行复杂推理的能力。
  • 闭源模型在直接给上下文时准确率约70%,但需从长文档检索时降至50%以下。
  • 开源模型最高仅27%准确率,适合研究长文档多模态理解的挑战。

文档是保存和传播信息的基础,常包含复杂版式、表格和图表,给自动文档理解(DU)带来巨大挑战。尽管视觉语言大模型(VLLMs)在多项任务中表现提升,其处理长上下文视觉输入的有效性仍不明确。本文提出WikiMixQA,一个包含1,000道多选题的多模态基准,用于评估从4,000个维基百科页面提取的表格和图表上的跨模态推理能力,覆盖七个不同主题。与现有基准不同,WikiMixQA强调复杂推理,要求模型综合多个模态信息。我们评估了12个最先进的视觉语言模型,发现虽然专有模型在提供直接上下文时准确率达约70%,但在需要从长文档中检索时性能显著下降;其中仅GPT-4-o在该设置下准确率超过50%,而开源模型表现更差,最高仅为27%。这些结果凸显了长上下文多模态推理的挑战,并确立了WikiMixQA作为推动文档理解研究的关键基准。

原文摘要 · Abstract (English)

Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU). While vision-language large models (VLLMs) have demonstrated improvements across various tasks, their effectiveness in processing long-context vision inputs remains unclear. This paper introduces WikiMixQA, a benchmark comprising 1,000 multiple-choice questions (MCQs) designed to evaluate cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages spanning seven distinct topics. Unlike existing benchmarks, WikiMixQA emphasizes complex reasoning by requiring models to synthesize information from multiple modalities. We evaluate 12 state-of-the-art vision-language models, revealing that while proprietary models achieve ~70% accuracy when provided with direct context, their performance deteriorates significantly when retrieval from long documents is required. Among these, GPT-4-o is the only model exceeding 50% accuracy in this setting, whereas open-source models perform considerably worse, with a maximum accuracy of 27%. These findings underscore the challenges of long-context, multi-modal reasoning and establish WikiMixQA as a crucial benchmark for advancing document understanding research.

多模态文档理解推理评测表格分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。