用高频感知与阴影掩膜指导,精准恢复文档阴影下的细节。
MatteViT: High-Frequency-Aware Document Shadow Removal with Shadow Matte Guidance
- 引入高频增强模块和连续亮度阴影掩膜,双策略保细节。
- 在RDD和Kligler数据集上达到当前最佳性能,提升文字识别效果。
- 适合需要高精度文档修复的OCR应用,尤其关注文字边缘保真。
文档阴影去除对提升数字化文档清晰度至关重要。保留高频细节(如文字边缘和线条)尤为关键,因阴影常会遮蔽或扭曲精细结构。本文提出一种新型阴影去除框架MatteViT,通过结合空间与频域信息,在消除阴影的同时保持细粒度结构。为有效保留这些细节,提出两种策略:首先,设计轻量级高频放大模块(HFAM),自适应分解并增强高频成分;其次,构建基于连续亮度的阴影掩膜,利用自建掩膜数据集与生成器,在早期处理阶段提供精确空间引导。该方法能准确识别细粒度区域,并高保真恢复。在公开基准数据集RDD和Kligler上的大量实验表明,MatteViT达到当前最优性能,为真实场景文档去影提供了鲁棒且实用的解决方案。此外,该方法在下游任务(如光学字符识别)中更好保留文本级细节,显著提升识别准确率。
原文摘要 · Abstract (English)
Document shadow removal is essential for enhancing the clarity of digitized documents. Preserving high-frequency details (e.g., text edges and lines) is critical in this process because shadows often obscure or distort fine structures. This paper proposes a matte vision transformer (MatteViT), a novel shadow removal framework that applies spatial and frequency-domain information to eliminate shadows while preserving fine-grained structural details. To effectively retain these details, we employ two preservation strategies. First, our method introduces a lightweight high-frequency amplification module (HFAM) that decomposes and adaptively amplifies high-frequency components. Second, we present a continuous luminance-based shadow matte, generated using a custom-built matte dataset and shadow matte generator, which provides precise spatial guidance from the earliest processing stage. These strategies enable the model to accurately identify fine-grained regions and restore them with high fidelity. Extensive experiments on public benchmarks (RDD and Kligler) demonstrate that MatteViT achieves state-of-the-art performance, providing a robust and practical solution for real-world document shadow removal. Furthermore, the proposed method better preserves text-level details in downstream tasks, such as optical character recognition, improving recognition performance over prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。