arXiv:2505.05446cs.CVcs.CL2025-05CVPR被引 7

用自适应标记语言生成提升文档理解的上下文准确性

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

  • 通过生成Markdown、JSON等标记语言构建结构化文档表示
  • 在380万条预训练数据和62.4万条指令数据上表现优于现有模型
  • 适合需要精准理解复杂布局文档的研究者与开发者

随着文本密集型视觉内容增多,视觉文档理解变得愈发重要。该领域面临视觉感知与文本理解有效融合的挑战,尤其在多样化的复杂版式文档中。现有微调数据集常缺乏足够上下文信息,导致幻觉和空间关系理解不足。为此,我们提出一种创新流程,利用自适应生成标记语言(如Markdown、JSON、HTML、TiKZ)构建高度结构化的文档表征,并实现上下文扎根的响应。我们引入两个细粒度结构化数据集:包含约380万条预训练数据对的DocMark-Pile,以及包含62.4万条微调标注的DocMark-Instruct。大量实验表明,所提模型在多个视觉文档理解基准上显著优于现有先进多模态大模型,提升了复杂视觉场景下的高级推理与理解能力。代码与模型已开源至https://github.com/Euphoria16/DocMark。

原文摘要 · Abstract (English)

Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex layouts. Moreover, existing fine-tuning datasets for this domain often fall short in providing the detailed contextual information for robust understanding, leading to hallucinations and limited comprehension of spatial relationships among visual elements. To address these challenges, we propose an innovative pipeline that utilizes adaptive generation of markup languages, such as Markdown, JSON, HTML, and TiKZ, to build highly structured document representations and deliver contextually-grounded responses. We introduce two fine-grained structured datasets: DocMark-Pile, comprising approximately 3.8M pretraining data pairs for document parsing, and DocMark-Instruct, featuring 624k fine-tuning data annotations for grounded instruction following. Extensive experiments demonstrate that our proposed model significantly outperforms existing state-of-theart MLLMs across a range of visual document understanding benchmarks, facilitating advanced reasoning and comprehension capabilities in complex visual scenarios. Our code and models are released at https://github. com/Euphoria16/DocMark.

文档理解标记生成多模态结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。