arXiv:2505.21975cs.CV2025-05SIGGRAPH被引 3

用坐标扩散模型实现文档去畸变,更好保留结构细节。

DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model

  • 以坐标级去噪替代像素级,生成变形矫正映射
  • 在6300张真实图像上达到最优性能,结构保留更佳
  • 适合需要高保真文档修复的研究与应用

文档去畸变旨在校正拍摄文档图像中的形变,提升文本可读性,虽已取得显著进展,但仍难以有效保持文档结构。受扩散模型近期发展的启发,本文提出DvD,首个基于扩散框架的文档去畸变生成模型。不同于传统像素级去噪,DvD采用坐标级去噪,生成用于形变矫正的映射。同时,引入时变条件精炼机制,增强文档结构的保持能力。实验发现现有基准无法全面评估去畸变模型,因此构建了AnyPhotoDoc6300——一个包含6300对真实图像、覆盖三个不同领域的大型基准,支持细粒度评估。大量实验证明,DvD在多个基准(包括DocUNet、DIR300和AnyPhotoDoc6300)上均达到领先性能,且计算效率合理。代码与数据集将公开于https://github.com/hanquansanren/DvD。

原文摘要 · Abstract (English)

Document dewarping aims to rectify deformations in photographic document images, thus improving text readability, which has attracted much attention and made great progress, but it is still challenging to preserve document structures. Given recent advances in diffusion models, it is natural for us to consider their potential applicability to document dewarping. However, it is far from straightforward to adopt diffusion models in document dewarping due to their unfaithful control on highly complex document images (e.g., 2000$times$3000 resolution). In this paper, we propose DvD, the first generative model to tackle document Dewarping via a Diffusion framework. To be specific, DvD introduces a coordinate-level denoising instead of typical pixel-level denoising, generating a mapping for deformation rectification. In addition, we further propose a time-variant condition refinement mechanism to enhance the preservation of document structures. In experiments, we find that current document dewarping benchmarks can not evaluate dewarping models comprehensively. To this end, we present AnyPhotoDoc6300, a rigorously designed large-scale document dewarping benchmark comprising 6,300 real image pairs across three distinct domains, enabling fine-grained evaluation of dewarping models. Comprehensive experiments demonstrate that our proposed DvD can achieve state-of-the-art performance with acceptable computational efficiency on multiple metrics across various benchmarks, including DocUNet, DIR300, and AnyPhotoDoc6300. The new benchmark and code will be publicly available at https://github.com/hanquansanren/DvD.

文档去畸变扩散模型坐标生成结构保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。