用在线强化学习提升盲脸修复质量,减少失真和身份模糊。
Enhancing Blind Face Restoration through Online Reinforcement Learning
- 引入在线强化学习框架,动态优化修复策略
- 在多个数据集上实现优于基线的修复效果
- 适合需要高保真人脸修复的应用场景
盲脸修复(BFR)面临解空间过大导致的常见问题,如细节丢失和身份混淆。为此,我们提出首个应用于BFR任务的在线强化学习框架——似然正则化策略优化(LRPO)。LRPO通过采样候选结果获得奖励,优化策略网络,在提高高质量输出概率的同时增强对低质量输入的修复能力。为解决直接应用强化学习导致的结果偏离真实值的问题,我们提出三项关键策略:1)专为面部修复设计的复合奖励函数;2)基于真实图像的似然正则化;3)按噪声水平分配优势值。大量实验表明,所提方法显著优于基线模型,在多个数据集上达到当前最优性能。
原文摘要 · Abstract (English)
Blind Face Restoration (BFR) encounters inherent challenges in exploring its large solution space, leading to common artifacts like missing details and identity ambiguity in the restored images. To tackle these challenges, we propose a Likelihood-Regularized Policy Optimization (LRPO) framework, the first to apply online reinforcement learning (RL) to the BFR task. LRPO leverages rewards from sampled candidates to refine the policy network, increasing the likelihood of high-quality outputs while improving restoration performance on low-quality inputs. However, directly applying RL to BFR creates incompatibility issues, producing restoration results that deviate significantly from the ground truth. To balance perceptual quality and fidelity, we propose three key strategies: 1) a composite reward function tailored for face restoration assessment, 2) ground-truth guided likelihood regularization, and 3) noise-level advantage assignment. Extensive experiments demonstrate that our proposed LRPO significantly improves the face restoration quality over baseline methods and achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。