arXiv:2603.09414cs.CVcs.AI2026-03中稿 · IEEE TMM

用描述性知识引导布局分析,让模型更懂不同文档的结构差异。

PromptDLA: A Domain-aware Prompt Document Layout Analysis Framework with Descriptive Knowledge as a Cue

  • 基于领域特征定制提示词,融入文档类型、语言等先验知识。
  • 在DocLayNet等四个数据集上达到当前最佳性能。
  • 适合需要跨领域通用布局识别的文档智能任务。

文档版面分析(DLA)对文档人工智能至关重要,近年来大规模公开数据集涌现。现有方法常将多领域数据混合训练以提升泛化能力,但直接合并会导致性能下降,因忽略了各领域固有的版面结构差异,如标注风格、文档类型和语言不同。本文提出PromptDLA,一种基于描述性知识的领域感知提示框架,通过定制化提示词引入领域先验。该提示器根据数据领域的具体属性生成提示,引导模型聚焦关键特征与结构,显著提升跨域泛化能力。大量实验表明,该方法在DocLayNet、PubLayNet、M6Doc和D$^4$LA四个基准上均取得当前最优表现。代码已开源:https://github.com/Zirui00/PromptDLA。

原文摘要 · Abstract (English)

Document Layout Analysis (DLA) is crucial for document artificial intelligence and has recently received increasing attention, resulting in an influx of large-scale public DLA datasets. Existing work often combines data from various domains in recent public DLA datasets to improve the generalization of DLA. However, directly merging these datasets for training often results in suboptimal model performance, as it overlooks the different layout structures inherent to various domains. These variations include different labeling styles, document types, and languages. This paper introduces PromptDLA, a domain-aware Prompter for Document Layout Analysis that effectively leverages descriptive knowledge as cues to integrate domain priors into DLA. The innovative PromptDLA features a unique domain-aware prompter that customizes prompts based on the specific attributes of the data domain. These prompts then serve as cues that direct the DLA toward critical features and structures within the data, enhancing the model's ability to generalize across varied domains. Extensive experiments show that our proposal achieves state-of-the-art performance among DocLayNet, PubLayNet, M6Doc, and D$^4$LA. Our code is available at https://github.com/Zirui00/PromptDLA.

文档分析提示工程领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。