arXiv:2512.20236cs.CV2025-12中稿 · ICDAR 2025被引 1

构建首个覆盖11种印地语系语言的文档布局数据集,解决多语言文档理解难题。

IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing

  • 构建涵盖11种印地语系语言和12类文档的大型标注数据集
  • 在新数据集上微调英文模型,性能显著提升
  • 适用于多语言文档数字化与跨语言布局分析研究

文档布局分析对信息检索、提取、OCR和数字化等下游任务至关重要。然而,现有大规模数据集如PubLayNet和DocBank缺乏细粒度区域标签和多语言多样性,难以刻画复杂文档结构。人类标注数据集如M6Doc和D4LA虽标签更丰富、领域更广,但规模小且多语言覆盖不足。这一缺口在印地语系文档中尤为突出——其包含多种书写系统却在现有数据集中严重缺失,制约了该领域的进展。为此,我们提出IndicDLP,一个涵盖11种代表性印地语系语言及英语、覆盖12类常见文档领域的大型基础文档布局数据集。此外,我们从DocLayNet和M6Doc中构建了UED-mini数据集,用于增强预训练并为印地语布局模型提供坚实基础。实验表明,在IndicDLP上微调现有英文模型可显著提升性能,且模型在印地语布局外也具备良好泛化能力,验证了其有效性。本工作填补了规模、多样性与标注精细度的空白,推动更具包容性与效率的文档理解发展。

原文摘要 · Abstract (English)

Document layout analysis is essential for downstream tasks such as information retrieval, extraction, OCR, and digitization. However, existing large-scale datasets like PubLayNet and DocBank lack fine-grained region labels and multilingual diversity, making them insufficient for representing complex document layouts. In contrast, human-annotated datasets such as M6Doc and D4LA offer richer labels and greater domain diversity, but are too small to train robust models and lack adequate multilingual coverage. This gap is especially pronounced for Indic documents, which encompass diverse scripts yet remain underrepresented in current datasets, further limiting progress in this space. To address these shortcomings, we introduce IndicDLP, a large-scale foundational document layout dataset spanning 11 representative Indic languages alongside English and 12 common document domains. Additionally, we curate UED-mini, a dataset derived from DocLayNet and M6Doc, to enhance pretraining and provide a solid foundation for Indic layout models. Our experiments demonstrate that fine-tuning existing English models on IndicDLP significantly boosts performance, validating its effectiveness. Moreover, models trained on IndicDLP generalize well beyond Indic layouts, making it a valuable resource for document digitization. This work bridges gaps in scale, diversity, and annotation granularity, driving inclusive and efficient document understanding.

文档布局多语言数据集印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。