arXiv:2605.01018cs.CV2026-05被引 3

首个面向真实场景表格图像理解的评测基准,揭示现有模型严重不足。

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

论文配图:WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild
图 1 · 摘自论文原文
  • 构建包含402张真实表格图像的问答数据集,覆盖17种问题类型。
  • 21个主流多模态模型平均准确率仅18.6%,最高不足50%。
  • 诊断发现结构感知与数值推理是模型核心短板,适合评估视觉理解能力。

利用多模态基础模型分析表格图像在消费和企业场景中具有高价值但极具挑战性。当前评估多依赖结构化文本或清晰渲染图像,忽视了真实场景表格图像的视觉复杂性。这些图像布局多样、领域广泛,需强结构感知与数值推理能力。为此,我们提出WildTableBench,首个针对真实世界表格图像的问答评测基准。该数据集包含402张从在线论坛和网站收集的高信息密度表格图像,涵盖多个领域;配套928个人工标注并验证的问题,分为5大类共17种子类型。我们评估了21个前沿开源及专有模型,仅有1个模型准确率超过50%,其余模型表现介于4.1%至49.9%之间。进一步诊断分析揭示了模型在结构感知与推理方面的持续弱点。结果与分析为当前模型能力提供了重要洞察,并确立WildTableBench作为表格图像理解的有效诊断基准。

原文摘要 · Abstract (English)

Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its importance, current evaluations rely largely on structured-text tables or clean rendered images, leaving the visual complexity of in-the-wild table images underexplored. Such images feature varied layouts and diverse domains that demand sophisticated structural perception and numerical reasoning. To bridge this gap, we introduce WildTableBench, the first question-answering benchmark for naturally occurring table images from real-world settings. WildTableBench comprises 402 high-information-density table images collected from online forums and websites across diverse domains, together with 928 manually annotated and verified questions spanning 17 subtypes across five categories. We evaluate 21 frontier proprietary and open-source multimodal foundation models on this benchmark. Only one model exceeds 50% accuracy, while all remaining models range from 4.1% to 49.9%. We further conduct diagnostic analyses to characterize model failures and reveal persistent weaknesses in structural perception and reasoning. These results and analyses provide useful insights into current model capabilities and establish WildTableBench as a valuable diagnostic benchmark for table image understanding. Dataset: https://huggingface.co/datasets/jzhuang/WildTableBench Code: https://github.com/hjzhe/WildTableBench Leaderboard: https://hjzhe.github.io/WildTableBench

表格理解多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。