arXiv:2603.23885cs.CV2026-03被引 5

用真实场景合成数据和结构感知训练,提升文档解析的准确性和鲁棒性

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

  • 通过合成真实文档场景生成大规模标注数据
  • 在真实拍摄文档上达到更高解析准确率
  • 适合需要强鲁棒性的文档理解研究者

文档解析近年借助多模态大模型直接从图像映射到结构化输出取得进展。传统级联方法依赖精确版面分析,在非标准或随意拍摄场景下易失效。尽管端到端方法缓解了这一依赖,仍存在重复、幻觉和结构不一致问题,主要源于缺乏大规模高质量的全页端到端标注数据以及缺乏结构感知训练策略。为此,我们提出一种数据与训练协同设计框架,以实现鲁棒的端到端文档解析。通过真实场景合成策略,将版面模板与丰富文档元素组合,构建大规模、结构多样的全页监督数据;通过文档感知训练方案,引入渐进式学习和结构标记优化,增强结构保真度与解码稳定性。我们还构建了基于真实拍摄文档的Wild-OmniDocBench基准用于鲁棒性评估。集成至10亿参数的MLLM后,该方法在扫描/数字文档及真实拍摄场景中均表现更优。所有模型、数据合成流程与基准将公开发布,推动文档理解研究发展。

原文摘要 · Abstract (English)

Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend on precise layout analysis and often fail under casually captured or non-standard conditions. Although end-to-end approaches mitigate this dependency, they still exhibit repetitive, hallucinated, and structurally inconsistent predictions - primarily due to the scarcity of large-scale, high-quality full-page (document-level) end-to-end parsing data and the lack of structure-aware training strategies. To address these challenges, we propose a data-training co-design framework for robust end-to-end document parsing. A Realistic Scene Synthesis strategy constructs large-scale, structurally diverse full-page end-to-end supervision by composing layout templates with rich document elements, while a Document-Aware Training Recipe introduces progressive learning and structure-token optimization to enhance structural fidelity and decoding stability. We further build Wild-OmniDocBench, a benchmark derived from real-world captured documents for robustness evaluation. Integrated into a 1B-parameter MLLM, our method achieves superior accuracy and robustness across both scanned/digital and real-world captured scenarios. All models, data synthesis pipelines, and benchmarks will be publicly released to advance future research in document understanding.

文档解析多模态合成数据结构感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。