arXiv:2501.05497cs.CLcs.AI2025-01

用合成数据提升小模型文档布局生成与分类能力

Spatial Information Integration in Small Language Models for Document Layout Generation and Classification

  • 通过生成合成布局数据弥补公开半结构化文档数据不足
  • 新方法在布局生成上优于LayoutTransformer,提升文本分类准确率
  • 适合需要小模型高效处理文档结构的开发者使用

文档布局理解旨在分析文档中信息的空间排列以解析其结构。现有模型如LayoutLM及其后续版本在半结构化文档上表现优异,但缺乏公开的半结构化数据集成为主要瓶颈。尽管日常中常见如对账单、采购订单、收据等半结构化文档,却缺少可用于训练机器学习模型的公开数据集。本文提出一种生成新型合成布局信息的方法,以缓解数据短缺问题。实验结果表明,该方法在布局生成性能上优于LayoutTransformer;同时,在某些场景下,结合边界框信息可提升文本分类效果。

原文摘要 · Abstract (English)

Document layout understanding is a field of study that analyzes the spatial arrangement of information in a document hoping to understand its structure and layout. Models such as LayoutLM (and its subsequent iterations) can understand semi-structured documents with SotA results; however, the lack of open semi-structured data is a limitation in itself. While semi-structured data is common in everyday life (balance sheets, purchase orders, receipts), there is a lack of public datasets for training machine learning models for this type of document. In this investigation we propose a method to generate new, synthetic, layout information that can help overcoming this data shortage. According to our results, the proposed method performs better than LayoutTransformer, another popular layout generation method. We also show that, in some scenarios, text classification can improve when supported by bounding box information.

文档布局小模型合成数据布局生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。