通过自强化奖励机制,让无监督去雨模型自动筛选高质量结果并提升性能。
Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategy

- 利用图像质量评估动态挑选训练中的优质去雨结果作为奖励信号。
- 在多个数据集上达到当前最优,主观与客观指标均领先现有方法。
- 策略可适配其他方法,且能兼容已有有监督网络,泛化性强。
无监督去雨因无需成对标注而受到关注,但缺乏强约束导致网络难收敛,尤其面对雨渍多样性的复杂退化。关键思路是:训练中偶尔会出现高质量去雨结果,可用来引导优化。为此,提出RGSUD(基于奖励引导的自强化无监督图像去雨),包含两个阶段:奖励循环与自强化(SR)训练。第一阶段设计基于图像质量评估(IQA)的动态奖励循环机制,从训练中选出最优去雨输出,并持续收集高质量去雨图像。第二阶段将这些奖励融入模型优化过程,缩小优化空间,增强去雨输出与真实干净图像的一致性。通过IQA引导的自强化损失和动态更新的奖励,提升伪成对数据合成质量并稳定优化过程。大量实验表明,该方法在多个数据集(包括合成成对、真实成对和真实无配对图像)上均达当前最优,主观与客观评价指标均优于现有无监督去雨方法。此外,自强化策略可适配其他无监督去雨方法,所提框架对已有有监督去雨网络具有强泛化能力。
原文摘要 · Abstract (English)
Unsupervised deraining has attracted attention for its ability to learn the real-world distribution of rain without paired supervision. However, the lack of strong constraints makes it difficult for the network to converge, especially with the complex diversity of rain degradation. A key motivation is that high-quality deraining results occasionally emerge during training, which can be leveraged to guide the optimization process. To overcome these challenges, we introduce RGSUD (Reward-Guided Self-Reinforcement Unsupervised Image Deraining), comprising two key stages: reward recycling and self-reinforcement (SR) training. For the former stage, we propose an Image Quality Assessment (IQA)-based dynamic reward recycling mechanism that selects optimal derained outputs during training and continuously collects high-quality deraining images. In latter stage, we incorporate these rewards into the model's optimization process, constraining the optimization space and improving alignment between derained outputs and clean images. By leveraging IQA-based self-reinforced loss and dynamically updated rewards, we enhance the quality of synthesized pseudo-paired data and stabilize the optimization. Extensive experiments demonstrate that our method achieves SOTA performance across multiple datasets, including paired synthetic, paired real, and unpaired real images, outperforming existing unsupervised deraining approaches in both subjective and objective IQA metrics. Additionally, we show that the self-reinforcement strategy is adaptable to other unsupervised deraining methods and our deraining framework demonstrates strong generalization across existing supervised deraining networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。