构建400万例真实表格数据集,推动视觉表理解模型更鲁棒。
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- 基于200万唯一真实表格,保留88%原始可视化
- 包含21项任务,支持跨任务泛化与未见场景测试
- 适合训练和评估复杂视觉表理解模型
尽管表格理解日益依赖纯像素输入,现有基准多采用合成渲染数据,缺乏真实世界表格的复杂性与视觉多样性。此外,现有视觉表理解(VTU)数据集仅提供固定样例与单一可视化,无法获取底层序列化数据以支持重表述。我们提出TABLET,一个大规模VTU数据集,包含400万例样本,覆盖21个任务,源自200万唯一表格,其中88%保留原始可视化。为评估模型对表格与视觉内容的联合推理能力,我们还引入VisualTableQA,要求同时具备视觉感知与表格理解能力。在TABLET上微调Qwen2.5-VL-7B和Gemma 3-4B等视觉语言模型,可提升其在已见与未见任务上的表现,并增强对真实表格图像的鲁棒性。通过保留原始可视化并实现示例可追溯性,TABLET为未来VTU模型的鲁棒训练与可扩展评估奠定基础。
原文摘要 · Abstract (English)
While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of real-world tables. Additionally, existing visual table understanding (VTU) datasets offer fixed examples with single visualizations and pre-defined instructions, providing no access to underlying serialized data for reformulation. We introduce TABLET, a large-scale VTU dataset with 4 million examples across 21 tasks, grounded in 2 million unique tables where 88% preserve original visualizations. To evaluate whether models are able to jointly reason over tabular and visual content, we also introduce VisualTableQA, a benchmark requiring both visual perception and table understanding. Fine-tuning vision-language models like Qwen2.5-VL-7B and Gemma 3-4B on TABLET improves performance on seen and unseen VTU tasks while increasing robustness on real-world table visualizations. By preserving original visualizations and maintaining example traceability in a unified large-scale collection, TABLET establishes a foundation for robust training and extensible evaluation of future VTU models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。