arXiv:2504.14238cs.CV2025-04

构建首个高分辨率文档反光去除数据集,提出无标注定位先验网络。

Single Document Image Highlight Removal via A Large-Scale Real-World Dataset and A Location-Aware Network

  • 基于残差图设计无标注反光位置先验,自动估计反光区域。
  • 在14,902张真实图像上实现最优性能,PSNR提升5.01%,RMSE下降13.17%。
  • 适合图像修复、文档数字化及低光照场景下视觉增强研究者使用。

反射性文档在环境光下常出现镜面反光,严重干扰文字可读性并降低整体视觉质量。尽管深度学习方法在反光去除方面展现出潜力,但针对文档图像仍不理想,主要受限于缺乏专用数据集和适配的网络结构。为此,我们构建了DocHR14K,一个包含14,902对高分辨率图像的大规模真实世界数据集,覆盖六类文档及多种光照条件。据我们所知,这是首个捕捉广泛真实光照条件的高分辨率文档反光去除数据集。受残差图能自然揭示反光区域空间结构的启发,我们提出一种无需人工标注的反光位置先验(HLP),用于估计反光掩码。在此基础上,提出位置感知的拉普拉斯金字塔反光去除网络(L2HRNet),通过融合估计先验与扩散模块有效去除反光并恢复细节。大量实验表明,DocHR14K显著提升了复杂光照下的反光去除效果。我们的L2HRNet在三个基准数据集上达到当前最佳性能,尤其在DocHR14K上实现PSNR提升5.01%,RMSE降低13.17%。

原文摘要 · Abstract (English)

Reflective documents often suffer from specular highlights under ambient lighting, severely hindering text readability and degrading overall visual quality. Although recent deep learning methods show promise in highlight removal, they remain suboptimal for document images, primarily due to the lack of dedicated datasets and tailored architectural designs. To tackle these challenges, we present DocHR14K, a large-scale real-world dataset comprising 14,902 high-resolution image pairs across six document categories and various lighting conditions. To the best of our knowledge, this is the first high-resolution dataset for document highlight removal that captures a wide range of real-world lighting conditions. Additionally, motivated by the observation that the residual map between highlighted and clean images naturally reveals the spatial structure of highlight regions, we propose a simple yet effective Highlight Location Prior (HLP) to estimate highlight masks without human annotations. Building on this prior, we present the Location-Aware Laplacian Pyramid Highlight Removal Network (L2HRNet), which effectively removes highlights by leveraging estimated priors and incorporates diffusion module to restore details. Extensive experiments demonstrate that DocHR14K improves highlight removal under diverse lighting conditions. Our L2HRNet achieves state-of-the-art performance across three benchmark datasets, including a 5.01\% increase in PSNR and a 13.17\% reduction in RMSE on DocHR14K.

图像修复文档处理反光去除无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。