用AI自动生成海量高质量表格数据,提升表格识别模型性能。
TableNet A Large-Scale Table Dataset with LLM-Powered Autonomous

- 构建首个由大模型驱动的自主表格生成与识别系统。
- 生成百万级表格图像,识别准确率显著优于现有方法。
- 适合需要高质量表格数据的AI研究者和开发者使用。
表格结构识别(TSR)需依赖大语言模型(LLM)的逻辑推理能力以处理复杂布局,但现有数据集规模与质量有限,难以充分发挥该能力。为此,我们提出TableNet数据集,通过多源收集与生成构建。核心是首个由大模型驱动的自主表格生成与识别多智能体系统。生成部分将可控的视觉、结构与语义参数融合于表格图像合成中,支持按用户配置生成语义一致、带标注的大规模表格,实现全面细致的标注分类体系。相比传统方法,该系统可理论无限生成、跨领域且风格灵活的表格图像,兼顾效率与精度。识别部分采用基于多样性的主动学习范式,整合多源表格,选择最具信息量的数据微调模型,在TableNet测试集上表现优异,同时训练样本大幅减少;在真实网络爬取表格上的性能远超基于主流数据集训练的模型。据我们所知,这是首个将主动学习应用于行/列数量、合并单元格、单元格内容等多样性特征丰富的表格结构识别的工作。
原文摘要 · Abstract (English)
Table Structure Recognition (TSR) requires the logical reasoning ability of large language models (LLMs) to handle complex table layouts, but current datasets are limited in scale and quality, hindering effective use of this reasoning capacity. We thus present TableNet dataset, a new table structure recognition dataset collected and generated through multiple sources. Central to our approach is the first LLM-powered autonomous table generation and recognition multi-agent system that we developed. The generation part of our system integrates controllable visual, structural, and semantic parameters into the synthesis of table images. It facilitates the creation of a wide array of semantically coherent tables, adaptable to user-defined configurations along with annotations, thereby supporting large-scale and detailed dataset construction. This capability enables a comprehensive and nuanced table image annotation taxonomy, potentially advancing research in table-related domains. In contrast to traditional data collection methods, This approach facilitates the theoretically infinite, domain-agnostic, and style-flexible generation of table images, ensuring both efficiency and precision. The recognition part of our system is a diversity-based active learning paradigm that utilizes tables from multiple sources and selectively samples most informative data to finetune a model, achieving a competitive performance on TableNet test set while reducing training samples by a large margin compared with baselines, and a much higher performance on web-crawled real-world tables compared with models trained on predominant table datasets. To the best of our knowledge, this is the first work which employs active learning into the structure recognition of tables which is diverse in numbers of rows or columns, merged cells, cell contents, etc, which fits better for diversity-based active learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。