构建合成金融表格数据集,提升表格信息提取精度。
SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction
- 通过合成方法生成带标注的金融表格数据
- 在真实金融表上测试模型,准确率优于现有生成模型
- 适合表格信息抽取与金融文档分析研究者使用
从文档图像中提取表格是具有挑战性的AI任务,而许多内容领域的标注数据难以获取。现有表格抽取数据集多聚焦于科学表格,因学术文章和源码资源丰富,但科学、金融等领域的表格在版式和排版上存在显著差异。当前数据集常缺乏表格内文字及其位置信息,依赖不可靠的OCR提取特征,影响自然语言处理模型训练效果。为此,我们提出SynFinTabs,一个大规模、带标注的合成金融表格数据集。我们的合成方法具备跨领域可迁移性。为验证数据集有效性,我们构建了基于布局的大语言模型FinTabQA,用于抽取式问答任务。在真实金融表格上的测试表明,该模型性能优于当前最先进的生成模型,结果具有可比性。我们已公开数据集、模型及生成代码。
原文摘要 · Abstract (English)
Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast amount of academic articles that are readily available, along with their source code. However, there are significant layout and typographical differences between tables found across scientific, financial, and other domains. Current datasets often lack the words, and their positions, contained within the tables, instead relying on unreliable OCR to extract these features for training modern machine learning models on natural language processing tasks. Therefore, there is a need for a more general method of obtaining labelled data. We present SynFinTabs, a large-scale, labelled dataset of synthetic financial tables. Our hope is that our method of generating these synthetic tables is transferable to other domains. To demonstrate the effectiveness of our dataset in training models to extract information from table images, we create FinTabQA, a layout large language model trained on an extractive question-answering task. We test our model using real-world financial tables and compare it to a state-of-the-art generative model and discuss the results. We make the dataset, model, and dataset generation code publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。