arXiv:2411.14957cs.CL2024-11中稿 · WACV 2025被引 1

无需真实标签,用合成标注训练模型提取多格式发票信息。

Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation

  • 用指令生成合成标签,不依赖真实标注数据。
  • 在内部票据上性能超Claude 3 Sonnet,速度快5倍、成本低85%。
  • 擅长识别罕见格式,适合企业反欺诈与报销审核场景。

员工提交的发票和收据是包含文本、视觉与版式信息的视觉丰富文档(VRDs)。为防范欺诈与滥用风险,组织需高效提取关键信息,用于评估费用合理性、合规性、票据有效性及异常检测。这些文档格式多样、语言各异、图像质量参差,且常缺乏真实标注,难以训练模型。本文提出任务感知指令标注(TAIL)方法,在无标签的VRD语料中生成合成标签,并通过基于响应的知识蒸馏微调多模态视觉丰富文档理解模型(VRDU),无需教师模型权重或训练数据即可生成适配格式的标注。在含真实标签的外部基准数据集上,本方法性能与Claude 3 Sonnet相当。在大型跨国企业内部票据上,该模型性能优于或持平于当前最佳大模态模型Claude 3 Sonnet,成本降低85%,速度提升约5倍,且在平均归一化莱文斯坦相似度(ANLS)上超越布局感知基线超过10%,得益于对罕见格式的信息推理能力。最后,展示了其在防止过度支付中的实际应用价值。

原文摘要 · Abstract (English)

Invoices and receipts submitted by employees are visually rich documents (VRDs) with textual, visual and layout information. To protect against the risk of fraud and abuse, it is crucial for organizations to efficiently extract desired information from submitted receipts. This helps in the assessment of key factors such as appropriateness of the expense claim, adherence to spending and transaction policies, the validity of the receipt, as well as downstream anomaly detection at various levels. These documents are heterogeneous, with multiple formats and languages, uploaded with different image qualities, and often do not contain ground truth labels for the efficient training of models. In this paper we propose Task Aware Instruction-based Labelling (TAIL), a method for synthetic label generation in VRD corpuses without labels, and fine-tune a multimodal Visually Rich Document Understanding Model (VRDU) on TAIL labels using response-based knowledge distillation without using the teacher model's weights or training dataset to conditionally generate annotations in the appropriate format. Using a benchmark external dataset where ground truth labels are available, we demonstrate conditions under which our approach performs at par with Claude 3 Sonnet through empirical studies. We then show that the resulting model performs at par or better on the internal expense documents of a large multinational organization than state-of-the-art LMM (large multimodal model) Claude 3 Sonnet while being 85% less costly and ~5X faster, and outperforms layout-aware baselines by more than 10% in Average Normalized Levenshtein Similarity (ANLS) scores due to its ability to reason and extract information from rare formats. Finally, we illustrate the usage of our approach in overpayment prevention.

文档理解合成数据知识蒸馏企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。