arXiv:2505.21771cs.CVcs.AI2025-05ACL被引 10

构建真实世界多模态表格评测集,揭示当前模型在视觉对齐与推理上的巨大短板。

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding

  • 人工构建500个真实场景多模态表格,含4021个问答对
  • 模型在视觉定位、空间对齐等任务上性能下降20%-40%
  • 适合研究多模态理解、表格推理与大模型评估的学者

多模态表格(即包含图表、地图、图标和颜色编码的表格)在真实应用中极为常见,但对多模态大语言模型(MLLMs)仍具挑战性。尽管文本与图像理解取得进展,针对以表格为中心的多模态推理的系统性评估仍有限。我们提出MMTABREAL,一个由人工标注的多模态表格基准,包含500个真实世界表格及4,021个问题-答案对。该数据集涵盖四种问题类型、五类推理类别和八种结构原型。对前沿模型的评估显示,在视觉定位、空间对齐和多步推理方面存在显著差距,性能较现有基准下降20%-40%。结果表明亟需更紧密融合视觉与表格结构的架构,并支持显式的数值/逻辑运算。MMTABREAL仅用于评估,提供一个反映真实多模态表格语言、结构与推理复杂性的严谨、可复现测试平台。

原文摘要 · Abstract (English)

Multimodal tables i.e. tabular layouts interleaved with charts, maps, icons, and color encodings are ubiquitous in real applications yet remain difficult for Multimodal Large Language Models (MLLMs). Despite advances in text and image understanding, systematic evaluation of table-centric multimodal reasoning is limited. We introduce MMTABREAL, a MultiModal Table Benchmark, human-curated suite of 500 real-world tables paired with 4,021 question-answer pairs. MMTABREAL spans four question types, five reasoning categories, and eight structural archetypes. Evaluations of state-of-the-art models reveal substantial gaps, especially in visual grounding, spatial alignment, and multi-step inference, with 20-40% performance drops relative to existing benchmarks. These results highlight the need for architectures that more tightly fuse vision with tabular structure and support explicit numeric/logical operations. MMTABREAL is released for evaluation only, providing a rigorous, reproducible testbed that reflects the linguistic, structural, and reasoning complexity of real-world multimodal tables.

多模态表格理解评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。