arXiv:2510.02133cs.AIcs.LG2025-10EMNLP被引 3

用参数化采样生成多语种合成文档,大幅降低标注成本

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

  • 通过概率建模布局与内容,可控生成多样文档
  • 提升关键信息提取准确率最高11%,标注量减少90%以上
  • 适合需要大规模低耗数据的工业级文档理解项目

企业级文档理解模型训练需大量多样且标注完善的多类型文档数据,但真实数据收集因隐私、法律限制及人工标注成本高昂,费用可达数百万美元。本文提出FlexDoc,一种可扩展的合成数据生成框架,结合随机模板与参数化采样,生成具有丰富标注的多语种半结构化文档。通过概率建模版式、视觉结构与内容差异性,实现规模化可控生成。在关键信息抽取(KIE)任务上的实验表明,使用其生成数据增强真实数据集,可使绝对F1分数提升最高11%,同时相比传统硬模板方法减少超过90%的标注工作量。该方案已在实际部署中应用,显著加速了企业级文档理解模型开发并大幅降低数据获取与标注成本。

原文摘要 · Abstract (English)

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to privacy constraints, legal restrictions, and the sheer volume of manual annotation needed - costs that can scale into millions of dollars. We introduce FlexDoc, a scalable synthetic data generation framework that combines Stochastic Schemas and Parameterized Sampling to produce realistic, multilingual semi-structured documents with rich annotations. By probabilistically modeling layout patterns, visual structure, and content variability, FlexDoc enables the controlled generation of diverse document variants at scale. Experiments on Key Information Extraction (KIE) tasks demonstrate that FlexDoc-generated data improves the absolute F1 Score by up to 11% when used to augment real datasets, while reducing annotation effort by over 90% compared to traditional hard-template methods. The solution is in active deployment, where it has accelerated the development of enterprise-grade document understanding models while significantly reducing data acquisition and annotation costs.

合成数据文档理解多语言降本增效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。