合成数据+双模型协同,提升专业文档信息提取精度
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
- 用结构化生成+领域查询构建高质量合成数据
- 通过迭代推理实现判别与生成模型联合优化
- 适合医疗金融等专业文档场景的低资源部署
领域特定的视觉丰富文档理解(VRDU)在医学、金融、材料科学等领域面临巨大挑战,因其文档复杂且敏感。现有大型多模态语言模型(LLMs/MLLMs)虽表现良好,但仍存在幻觉、领域适应不足及依赖大量微调数据集等问题。本文提出SynDoc框架,融合判别与生成模型以应对上述挑战。该框架采用稳健的合成数据生成流程,结合结构信息提取与领域特定查询生成,生成高质量标注数据。通过自适应指令微调,增强判别模型提取领域知识的能力;同时,利用递归推理机制迭代优化双模型输出,实现稳定准确的预测。该框架展现出可扩展、高效且精确的文档理解能力,弥合了领域特定适应性与通用世界知识之间的差距,适用于文档关键信息提取任务。
原文摘要 · Abstract (English)
Domain-specific Visually Rich Document Understanding (VRDU) presents significant challenges due to the complexity and sensitivity of documents in fields such as medicine, finance, and material science. Existing Large (Multimodal) Language Models (LLMs/MLLMs) achieve promising results but face limitations such as hallucinations, inadequate domain adaptation, and reliance on extensive fine-tuning datasets. This paper introduces SynDoc, a novel framework that combines discriminative and generative models to address these challenges. SynDoc employs a robust synthetic data generation workflow, using structural information extraction and domain-specific query generation to produce high-quality annotations. Through adaptive instruction tuning, SynDoc improves the discriminative model's ability to extract domain-specific knowledge. At the same time, a recursive inferencing mechanism iteratively refines the output of both models for stable and accurate predictions. This framework demonstrates scalable, efficient, and precise document understanding and bridges the gap between domain-specific adaptation and general world knowledge for document key information extraction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。