构建复杂表格空间推理基准,解决多模态模型感知过载问题。
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables
- 设计轨迹感知的基准数据集,显式关联逻辑推理与像素级空间位置。
- 在13个子类别上覆盖三类认知层级任务,含6799条高保真推理轨迹。
- 提出两阶段解耦框架,提升模型整体准确率12.3%,适合文档理解研究者。
结构化表格在金融、医疗和科研等专业领域是传递高密度信息的关键载体。尽管多模态大语言模型(MLLMs)取得进展,面对具有层次布局的复杂表格时,其推理性能仍受限。本文通过定量分析发现一个关键的感知瓶颈:随着任务复杂度提升,涉及的离散视觉区域数量不成比例增加,导致模型内部出现“感知过载”,难以维持隐式生成过程中的空间注意力。为此,我们提出TableVision——一个大规模、轨迹感知的基准,用于空间对齐的表格推理。TableVision将表格任务划分为感知、推理与分析三个认知层级,涵盖13个子类别。通过基于渲染的确定性对齐流程,数据集将多步逻辑推断与像素级空间真值明确绑定,包含6,799条高保真推理轨迹。实验结果表明,显式空间约束显著恢复了MLLM的推理能力。此外,我们的两阶段解耦框架在测试集上实现12.3%的整体准确率提升。TableVision为文档理解中的感知与逻辑协同提供了严格评测平台与新视角。
原文摘要 · Abstract (English)
Structured tables are essential for conveying high-density information in professional domains such as finance, healthcare, and scientific research. Despite the progress in Multimodal Large Language Models (MLLMs), reasoning performance remains limited for complex tables with hierarchical layouts. In this paper, we identify a critical Perception Bottleneck through quantitative analysis. We find that as task complexity scales, the number of involved discrete visual regions increases disproportionately. This processing density leads to an internal "Perceptual Overload," where MLLMs struggle to maintain accurate spatial attention during implicit generation. To address this bottleneck, we introduce TableVision, a large-scale, trajectory-aware benchmark designed for spatially grounded reasoning. TableVision stratifies tabular tasks into three cognitive levels (Perception, Reasoning, and Analysis) across 13 sub-categories. By utilizing a rendering-based deterministic grounding pipeline, the dataset explicitly couples multi-step logical deductions with pixel-perfect spatial ground truths, comprising 6,799 high-fidelity reasoning trajectories. Our empirical results, supported by diagnostic probing, demonstrate that explicit spatial constraints significantly recover the reasoning potential of MLLMs. Furthermore, our two-stage decoupled framework achieves a robust 12.3% overall accuracy improvement on the test set. TableVision provides a rigorous testbed and a fresh perspective on the synergy between perception and logic in document understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。