解决数码与拍摄文档解析中的结构错乱问题,提升复杂文档理解能力。
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

- 引入感知形变的训练机制,增强模型对拍摄文档几何畸变的鲁棒性。
- 在多个基准上达到96.87、88.53、78.41的最高分,超越现有方法。
- 适合需要高精度文档结构还原的科研、金融、法律等场景使用。
文档解析旨在将非结构化文档转化为结构化机器可读形式。尽管视觉语言模型(VLM)取得显著进展,现有方法仍面临两大挑战:其一,解耦式VLM方法严重依赖精确版面分析,而拍摄文档的几何畸变易引发级联误差;其二,端到端VLM方法虽减少对显式版面检测的依赖,但在高分辨率场景下常出现冗余生成、幻觉及结构推理不足。为此,本文提出NaviDC-OCR统一框架,引入形变感知学习以增强VLM的几何感知能力,并设计自适应采样机制表示复杂版面。此外,提出内容-结构解耦学习策略,显式建模公式语法与表格结构,实现更有效的结构化表征学习。大量实验表明,NaviDC-OCR在多个文档解析基准上表现领先,在OmniDocBench v1.6、Wild-OmniDocBench和PureDocBench上分别获得96.87、88.53和78.41的综合得分,并在ICDAR 2026 Sci-ImageMiner Challenge中排名第一,验证了其在复杂文档解析中的有效性与泛化能力。
原文摘要 · Abstract (English)
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。