构建首个大规模图文表格推理数据集,推动视觉语言模型理解复杂表格。
Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
- 用多模型协作生成数据,实现自动化、可扩展的高质量图文对构造。
- 包含2.5千张LaTeX渲染表格与6千个深度推理问答对,成本低于100美元。
- 适合研究视觉推理、多模态模型泛化能力的研究者使用。
在结构化数据(如表格)上的视觉推理是现代视觉-语言模型的关键能力,但现有基准在规模、多样性或推理深度方面仍显不足,尤其针对渲染后的表格图像。为此,我们提出Visual-TableQA,一个大规模、开放域的多模态数据集,专门用于评估和提升对复杂表格数据的视觉推理能力。我们的生成流程模块化、可扩展且完全自动化,涉及多个推理型大模型以不同角色协作:生成、验证与启发。Visual-TableQA包含2.5k个结构丰富的LaTeX渲染表格和6000个需要深度推理的问答对,总生成成本低于100美元。为增强多样性与创造力,流程通过跨模型提示('启发')实现多模型协同生成,并采用LLM裁判过滤机制。强模型引导布局与主题,弱模型加以拓展,共同提炼出多样化的推理模式与视觉结构。实验证明,在Visual-TableQA上微调的模型在外部基准上具有良好的泛化能力,表现优于若干专有模型,尽管数据为合成生成。完整流程与资源已公开于https://github.com/AI-4-Everyone/Visual-TableQA。
原文摘要 · Abstract (English)
Visual reasoning over structured data such as tables is a critical capability for modern vision-language models (VLMs), yet current benchmarks remain limited in scale, diversity, or reasoning depth, especially when it comes to rendered table images. Addressing this gap, we introduce Visual-TableQA, a large-scale, open-domain multimodal dataset specifically designed to evaluate and enhance visual reasoning over complex tabular data. Our generation pipeline is modular, scalable, and fully autonomous, involving multiple reasoning LLMs collaborating across distinct roles: generation, validation, and inspiration. Visual-TableQA comprises 2.5k richly structured LaTeX-rendered tables and 6k reasoning-intensive QA pairs, all produced at a cost of under USD 100. To promote diversity and creativity, our pipeline performs multi-model collaborative data generation via cross-model prompting ('inspiration') and LLM-jury filtering. Stronger models seed layouts and topics that weaker models elaborate, collectively distilling diverse reasoning patterns and visual structures into the dataset. Empirical results show that models fine-tuned on Visual-TableQA generalize robustly to external benchmarks, outperforming several proprietary models despite the dataset's synthetic nature. The full pipeline and resources are publicly available at https://github.com/AI-4-Everyone/Visual-TableQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。