arXiv:2506.11684cs.CVcs.AI2025-06EMNLP被引 10

评测大模型在图像表格间多跳推理的能力,填补视觉表格理解空白

MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space

  • 构建图像多表问答基准,要求跨表关联与多步推理
  • 3745个复杂问题显示主流模型表现严重不足
  • 提供指令微调数据集,显著提升模型视觉推理能力

视觉语言模型在解析视觉布局和文本方面表现出色,但在处理以图像形式呈现的多表格数据时仍存在鲁棒性与推理能力不足的问题,这在网页和数字文档中极为常见。现有基准大多仅涵盖单张表格或非视觉数据(如文本/结构化数据),未能评估模型对多样图像表格的解析能力、跨表信息关联及组合视觉数据上的多跳推理能力。为此,我们提出MTabVQA,一个专为多表格视觉问答设计的新基准。该基准包含3,745个需跨多个可视化表格进行多跳推理的复杂问题。我们对主流视觉语言模型在MTabVQA上的表现进行了全面评估,揭示其性能存在显著局限。进一步研究了后训练技术以增强推理能力,并发布了大规模指令微调数据集MTabVQA-Instruct。实验表明,使用MTabVQA-Instruct进行微调可显著提升模型在视觉多表推理任务上的表现。代码与数据集已公开于https://huggingface.co/datasets/mtabvqa/MTabVQA-Eval(匿名链接:https://anonymous.4open.science/r/MTabVQA-EMNLP-B16E)。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in interpreting visual layouts and text. However, a significant challenge remains in their ability to interpret robustly and reason over multi-tabular data presented as images, a common occurrence in real-world scenarios like web pages and digital documents. Existing benchmarks typically address single tables or non-visual data (text/structured). This leaves a critical gap: they don't assess the ability to parse diverse table images, correlate information across them, and perform multi-hop reasoning on the combined visual data. We introduce MTabVQA, a novel benchmark specifically designed for multi-tabular visual question answering to bridge that gap. MTabVQA comprises 3,745 complex question-answer pairs that necessitate multi-hop reasoning across several visually rendered table images. We provide extensive benchmark results for state-of-the-art VLMs on MTabVQA, revealing significant performance limitations. We further investigate post-training techniques to enhance these reasoning abilities and release MTabVQA-Instruct, a large-scale instruction-tuning dataset. Our experiments show that fine-tuning VLMs with MTabVQA-Instruct substantially improves their performance on visual multi-tabular reasoning. Code and dataset (https://huggingface.co/datasets/mtabvqa/MTabVQA-Eval) are available online (https://anonymous.4open.science/r/MTabVQA-EMNLP-B16E).

视觉问答多表推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。