针对多退化彩色文档图像,提出高效鲁棒的增强框架
GL-PGENet: A Parameterized Generation Framework for Robust Document Image Enhancement
- 分层架构融合全局校正与局部精修,实现由粗到细增强
- 参数化生成替代直接预测,提升局部一致性与泛化能力
- 适配真实场景,高分辨率下保持高效且跨域适应性强
文档图像增强(DIE)是文档AI系统的关键组件,其性能直接影响下游任务效果。针对现有方法仅限于单一退化修复或灰度图像处理的局限,本文提出全局与局部参数化生成增强网络(GL-PGENet),专为多退化彩色文档图像设计,兼顾效率与实际应用中的鲁棒性。核心创新包括:首先,采用分层增强框架,结合全局外观修正与局部精修,实现从粗到细的质量提升;其次,提出双分支局部精修网络,通过参数化生成机制替代传统像素级映射,利用学习到的中间参数表示生成输出,增强局部一致性并提升模型泛化能力;最后,改进NestUNet结构,引入密集块以有效融合低层像素特征与高层语义特征,适配文档图像特性。此外,采用两阶段训练策略:在包含50万+样本的合成数据集上大规模预训练,再进行任务特定微调。大量实验表明,GL-PGENet在DocUNet上取得0.7721的SSIM,在RealDAE上达到0.9480,均达当前最优水平,同时具备出色的跨域适应能力,高分辨率图像处理中无性能下降,证实其在真实场景中的实用性。
原文摘要 · Abstract (English)
Document Image Enhancement (DIE) serves as a critical component in Document AI systems, where its performance substantially determines the effectiveness of downstream tasks. To address the limitations of existing methods confined to single-degradation restoration or grayscale image processing, we present Global with Local Parametric Generation Enhancement Network (GL-PGENet), a novel architecture designed for multi-degraded color document images, ensuring both efficiency and robustness in real-world scenarios. Our solution incorporates three key innovations: First, a hierarchical enhancement framework that integrates global appearance correction with local refinement, enabling coarse-to-fine quality improvement. Second, a Dual-Branch Local-Refine Network with parametric generation mechanisms that replaces conventional direct prediction, producing enhanced outputs through learned intermediate parametric representations rather than pixel-wise mapping. This approach enhances local consistency while improving model generalization. Finally, a modified NestUNet architecture incorporating dense block to effectively fuse low-level pixel features and high-level semantic features, specifically adapted for document image characteristics. In addition, to enhance generalization performance, we adopt a two-stage training strategy: large-scale pretraining on a synthetic dataset of 500,000+ samples followed by task-specific fine-tuning. Extensive experiments demonstrate the superiority of GL-PGENet, achieving state-of-the-art SSIM scores of 0.7721 on DocUNet and 0.9480 on RealDAE. The model also exhibits remarkable cross-domain adaptability and maintains computational efficiency for high-resolution images without performance degradation, confirming its practical utility in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。