arXiv:2410.01609cs.CV2024-10中稿 · publication in Inf…被引 5

用合成数据+小样本校准,高效提取扫描文档关键信息

SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction

  • 用机器生成的合成数据做域适应,降低标注依赖
  • 在少量人工标注数据上校准,有效抑制噪声影响
  • 适合需要快速适配新领域的扫描文档信息提取任务

视觉丰富文档(VRDs)包含图表、表格和段落等元素,跨领域传递复杂信息。然而,从这些文档中提取关键信息仍需大量人力,尤其是对布局不一致且具有领域特性的扫描版文档。尽管预训练模型在VRD理解方面取得进展,但其对大规模标注数据的依赖限制了可扩展性。本文提出SynJAC(合成数据驱动的联合粒度自适应与校准),用于扫描文档中的关键信息提取。SynJAC利用机器生成的合成数据进行域适应,并在少量人工标注数据上进行校准以缓解噪声问题。通过融合细粒度与粗粒度文档表征学习,显著减少对大规模手动标注的需求,同时实现竞争力性能。大量实验验证了其在特定领域和扫描型VRD场景下的有效性。

原文摘要 · Abstract (English)

Visually Rich Documents (VRDs), comprising elements such as charts, tables, and paragraphs, convey complex information across diverse domains. However, extracting key information from these documents remains labour-intensive, particularly for scanned formats with inconsistent layouts and domain-specific requirements. Despite advances in pretrained models for VRD understanding, their dependence on large annotated datasets for fine-tuning hinders scalability. This paper proposes \textbf{SynJAC} (Synthetic-data-driven Joint-granular Adaptation and Calibration), a method for key information extraction in scanned documents. SynJAC leverages synthetic, machine-generated data for domain adaptation and employs calibration on a small, manually annotated dataset to mitigate noise. By integrating fine-grained and coarse-grained document representation learning, SynJAC significantly reduces the need for extensive manual labelling while achieving competitive performance. Extensive experiments demonstrate its effectiveness in domain-specific and scanned VRD scenarios.

文档理解合成数据少样本学习信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。