arXiv:2512.02498cs.CV2025-12被引 43

一个模型搞定多语言文档布局解析,端到端更准更快。

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

  • 用统一视觉语言模型同时完成布局识别、文字识别和关系理解。
  • 在126种语言的XDocParse上提升约10%性能,领先现有方法。
  • 适合需要跨语言文档智能处理的研究与工业应用。

文档布局解析是人工智能获取和理解全球结构化知识的关键入口,涵盖布局检测、文本识别和关系理解。当前方法依赖碎片化的多阶段流水线,存在误差传播问题,且难以实现联合训练优势。本文提出dots_ocr,首个在统一端到端框架中联合学习三项核心任务的视觉语言模型。其成功得益于一个可扩展的数据引擎,合成大规模多语言语料库,使模型在多种语言、布局和领域下表现稳健。在综合性基准OmniDocBench上达到当前最优性能;为进一步推动全球文档智能研究,我们发布XDocParse,覆盖126种语言的新基准。在该基准上,dots_ocr实现最先进的性能,相对提升约10%,展现出强大的多语言能力。

原文摘要 · Abstract (English)

Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recognition, and relational understanding, is particularly crucial for empowering next-generation Vision-Language Models. Current methods, however, rely on fragmented, multi-stage pipelines that suffer from error propagation and fail to leverage the synergies of joint training. In this paper, we introduce dots_ocr, a single Vision-Language Model that, for the first time, demonstrates the advantages of jointly learning three core tasks within a unified, end-to-end framework. This is made possible by a highly scalable data engine that synthesizes a vast multilingual corpus, empowering the model to deliver robust performance across a wide array of tasks, encompassing diverse languages, layouts, and domains. The efficacy of our unified paradigm is validated by state-of-the-art performance on the comprehensive OmniDocBench. Furthermore, to catalyze research in global document intelligence, we introduce XDocParse, a challenging new benchmark spanning 126 languages. On this benchmark, dots_ocr achieves state-of-the-art performance, delivering an approximately 10% relative improvement and demonstrating strong multilingual capability.

文档解析多语言视觉语言模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。