arXiv:2605.05955cs.CLcs.CV2026-05ACL

构建复杂表格多模态推理评测集,揭示当前模型在结构复杂时表现严重下降。

TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity

论文配图:TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
图 1 · 摘自论文原文
  • 设计多风格渲染管道生成3万张表格图像样本
  • 29个主流模型在复杂结构上性能显著下降
  • 适合研究表格理解、多模态鲁棒性的研究人员

我们提出TableVista,一个针对视觉与结构复杂性下的多模态表格推理的综合性评测基准。TableVista包含3,000个高质量表格推理问题,通过多风格渲染与变换管道,每个实例生成10种不同视觉变体,涵盖多样场景风格、鲁棒性扰动和仅视觉配置,共形成30,000个多模态样本,支持多维度评估。我们在TableVista上对29个最先进的开源与专有基础模型进行了广泛评估。通过全面的定量与定性分析发现,尽管模型在不同渲染风格下表现基本稳定,但在复杂结构布局和仅视觉设置下性能明显下降,表明当前模型在结构复杂性与视觉融合呈现结合时难以保持推理一致性。这些发现揭示了当前多模态能力的关键短板,为构建更鲁棒可靠的表格理解模型提供了重要启示。

原文摘要 · Abstract (English)

We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each instance is expanded into 10 distinct visual variants through our multi-style rendering and transformation pipeline. This process encompasses diverse scenario styles, robustness perturbations, and vision-only configurations, culminating in 30,000 multimodal samples for a multi-dimensional evaluation. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary foundation models on TableVista. Through comprehensive quantitative and qualitative analysis, we find that while evaluated models remain largely stable across diverse rendering styles, they exhibit pronounced performance degradation on complex structural layouts and vision-only settings, revealing that current models struggle to maintain reasoning consistency when structural complexity combines with visually integrated presentations. These findings highlight critical gaps in current multimodal capabilities, providing insights for advancing more robust and reliable table understanding models.

表格理解多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。