arXiv:2507.01422cs.CVcs.AI2025-07

用潜在空间扩散模型去除文档图像中的彩色阴影

DocShaDiffusion: Diffusion Model in Latent Space for Document Image Shadow Removal

  • 在潜在空间构建扩散模型,更高效捕捉文档特征
  • 提出软掩码生成模块,精准定位并处理彩色阴影区域
  • 自研大规模合成数据集,适合文档增强与图像修复研究

文档阴影去除是文档图像增强中的关键任务。现有方法多针对背景颜色恒定的阴影,忽视了彩色阴影问题。本文首次提出一种基于潜在空间的扩散模型DocShaDiffusion,将阴影图像从像素空间映射至潜在空间,提升特征提取效率。为应对彩色阴影,设计了阴影软掩码生成模块(SSGM),可生成精确掩码并针对性添加噪声。在此基础上,提出阴影掩码感知引导扩散模块(SMGDM),通过掩码指导扩散与去噪过程以实现阴影消除。此外,设计了一种抗阴影感知的感知损失函数,有效保留文档细节与结构。同时构建了一个大规模合成彩色阴影文档数据集(SDCSRD),模拟真实彩色阴影分布,为模型训练提供有力支持。在三个公开数据集上的实验验证了该方法优于当前最优方案。代码与数据集将开源。

原文摘要 · Abstract (English)

Document shadow removal is a crucial task in the field of document image enhancement. However, existing methods tend to remove shadows with constant color background and ignore color shadows. In this paper, we first design a diffusion model in latent space for document image shadow removal, called DocShaDiffusion. It translates shadow images from pixel space to latent space, enabling the model to more easily capture essential features. To address the issue of color shadows, we design a shadow soft-mask generation module (SSGM). It is able to produce accurate shadow mask and add noise into shadow regions specially. Guided by the shadow mask, a shadow mask-aware guided diffusion module (SMGDM) is proposed to remove shadows from document images by supervising the diffusion and denoising process. We also propose a shadow-robust perceptual feature loss to preserve details and structures in document images. Moreover, we develop a large-scale synthetic document color shadow removal dataset (SDCSRD). It simulates the distribution of realistic color shadows and provides powerful supports for the training of models. Experiments on three public datasets validate the proposed method's superiority over state-of-the-art. Our code and dataset will be publicly available.

文档增强扩散模型阴影去除图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。