arXiv:2507.19804cs.CV2025-07ICCV被引 4

聚焦文字区域,提升文档图像矫正精度

ForCenNet: Foreground-Centric Network for Document Image Rectification

  • 以前景元素为中心设计标签生成与掩码机制
  • 在四个真实场景数据集上达到最新最好效果
  • 适合需要高精度文档矫正的研究与应用

文档图像矫正旨在消除拍摄文档中的几何畸变,以促进文本识别。然而,现有方法常忽略前景元素的重要性,而这些元素为文档校正提供了关键的几何参考和版式信息。本文提出前景中心网络(ForCenNet),用于消除文档图像中的几何失真。具体而言,我们首先提出一种基于无畸变图像提取详细前景元素的前景中心标签生成方法;随后引入前景中心掩码机制,增强可读区域与背景之间的区分度;此外,设计曲率一致性损失,利用详细的前景标签帮助模型理解畸变分布。大量实验表明,ForCenNet在四个真实世界基准数据集(DocUNet、DIR300、WarpDoc、DocReal)上均达到新最优性能。定量分析显示,该方法能有效校正文本行、表格边框等版式元素。更多资源详见:https://github.com/caipeng328/ForCenNet。

原文摘要 · Abstract (English)

Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significance of foreground elements, which provide essential geometric references and layout information for document image correction. In this paper, we introduce Foreground-Centric Network (ForCenNet) to eliminate geometric distortions in document images. Specifically, we initially propose a foreground-centric label generation method, which extracts detailed foreground elements from an undistorted image. Then we introduce a foreground-centric mask mechanism to enhance the distinction between readable and background regions. Furthermore, we design a curvature consistency loss to leverage the detailed foreground labels to help the model understand the distorted geometric distribution. Extensive experiments demonstrate that ForCenNet achieves new state-of-the-art on four real-world benchmarks, such as DocUNet, DIR300, WarpDoc, and DocReal. Quantitative analysis shows that the proposed method effectively undistorts layout elements, such as text lines and table borders. The resources for further comparison are provided at https://github.com/caipeng328/ForCenNet.

文档矫正前景感知几何校正深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。