用大模型生成逼真发票,自动扩增训练数据。
Generating Synthetic Invoices via Layout-Preserving Content Replacement
- 用OCR提取原文本与布局,再以大模型生成新内容。
- 替换后通过修复技术保留原版排版与字体特征。
- 适合需要大量私有发票数据的文档智能研究者。
自动化发票处理的机器学习模型性能高度依赖大规模、多样化的数据集。然而,获取此类数据常受隐私法规限制且人工标注成本高昂。为此,我们提出一种新型管道,用于生成高保真的合成发票文档及其对应的结构化数据。该方法首先利用光学字符识别(OCR)提取源发票的文本内容与精确空间布局;随后,使用大语言模型(LLM)生成上下文合理的合成内容,替换部分数据字段;最后,采用图像修复技术消除原始文本,并以新内容渲染在原位置,完整保留布局和字体特征。该流程输出一对结果:一张视觉逼真的新发票图像,以及一个与合成内容完全对齐的结构化数据文件(JSON)。本方法为小规模、私有数据集提供可扩展、自动化的扩充方案,支持构建更大、更丰富的语料库,以训练更具鲁棒性与准确性的文档智能模型。
原文摘要 · Abstract (English)
The performance of machine learning models for automated invoice processing is critically dependent on large-scale, diverse datasets. However, the acquisition of such datasets is often constrained by privacy regulations and the high cost of manual annotation. To address this, we present a novel pipeline for generating high-fidelity, synthetic invoice documents and their corresponding structured data. Our method first utilizes Optical Character Recognition (OCR) to extract the text content and precise spatial layout from a source invoice. Select data fields are then replaced with contextually realistic, synthetic content generated by a large language model (LLM). Finally, we employ an inpainting technique to erase the original text from the image and render the new, synthetic text in its place, preserving the exact layout and font characteristics. This process yields a pair of outputs: a visually realistic new invoice image and a perfectly aligned structured data file (JSON) reflecting the synthetic content. Our approach provides a scalable and automated solution to amplify small, private datasets, enabling the creation of large, varied corpora for training more robust and accurate document intelligence models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。