arXiv:2511.23150cs.CV2025-11

分阶段逐步纠正任意文档图像的透视与物理扭曲,提升真实场景下文字识别效果。

Cascaded Robust Rectification for Arbitrary Document Images

  • 分三步逆向矫正:视角畸变→纸张弯曲→内容细节失真
  • 在多个基准上降低14.1%~34.7%的AAD指标,显著优于现有方法
  • 新评测指标可分离几何校正与OCR布局误差,适合边界不全文档

真实场景下的文档校正面临相机视角变化和物理形变的极大挑战。基于复杂变换可逐级分解的洞察,本文提出一种多阶段框架,以粗到精的方式逐步逆转不同类型的畸变。首先通过全局仿射变换校正由相机视角引起的透视失真;接着修正因纸张卷曲和折叠产生的几何形变;最后采用内容感知的迭代过程消除细粒度内容畸变。为克服现有评估协议的局限,我们还提出两种增强指标:布局对齐的OCR指标(AED/ACER),可将几何校正质量与OCR引擎的布局分析误差解耦;以及针对边界不完整的文档设计的掩码型AD/AAD(AD-M/AAD-M)。大量实验表明,该方法在多个挑战性基准上达到新的最优性能,使AAD指标下降14.1%–34.7%,在真实应用中表现优异。代码将公开于https://github.com/chaoyunwang/ArbDR。

原文摘要 · Abstract (English)

Document rectification in real-world scenarios poses significant challenges due to extreme variations in camera perspectives and physical distortions. Driven by the insight that complex transformations can be decomposed and resolved progressively, we introduce a novel multi-stage framework that progressively reverses distinct distortion types in a coarse-to-fine manner. Specifically, our framework first performs a global affine transformation to correct perspective distortions arising from the camera's viewpoint, then rectifies geometric deformations resulting from physical paper curling and folding, and finally employs a content-aware iterative process to eliminate fine-grained content distortions. To address limitations in existing evaluation protocols, we also propose two enhanced metrics: layout-aligned OCR metrics (AED/ACER) for a stable assessment that decouples geometric rectification quality from the layout analysis errors of OCR engines, and masked AD/AAD (AD-M/AAD-M) tailored for accurately evaluating geometric distortions in documents with incomplete boundaries. Extensive experiments show that our method establishes new state-of-the-art performance on multiple challenging benchmarks, yielding a substantial reduction of 14.1\%--34.7\% in the AAD metric and demonstrating superior efficacy in real-world applications. The code will be publicly available at https://github.com/chaoyunwang/ArbDR.

文档校正图像修复多阶段模型真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。