arXiv:2512.22647cs.CV2025-12中稿 · CVPR被引 3

提出细粒度感知奖励模型,解决超分辨率图像生成中的奖励欺骗问题。

FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution

论文配图:FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution
图 1 · 摘自论文原文
  • 设计基于编码器-解码器的细粒度感知奖励模型,同时输出全局质量分和局部失真图。
  • 在真实超分辨数据集FGR-30k上训练,能精准定位并量化细微伪影。
  • 引入协同进化课程学习机制,稳定训练过程,防止奖励误导,适合图像生成研究者。

基于人类反馈的强化学习(RLHF)在图像生成中通过奖励模型对齐人类偏好已证明有效。受此启发,将RLHF应用于图像超分辨率(ISR)任务,利用图像质量评估(IQA)模型作为奖励模型优化感知质量。然而,传统IQA模型通常仅输出单一全局评分,对局部细微失真不敏感,导致ISR模型可生成看似高分但感知劣质的伪影,造成优化目标与真实感知质量错位,引发奖励欺骗。为此,我们提出细粒度感知奖励模型(FinPercep-RM),采用编码器-解码器架构,在提供全局质量分数的同时,生成空间定位的感知退化图,精确量化局部缺陷。我们特别构建了包含多样且细微失真的真实超分辨数据集FGR-30k用于训练该模型。尽管FinPercep-RM性能优越,其复杂性带来生成策略学习的显著挑战,导致训练不稳定。为此,我们提出协同进化课程学习(CCL)机制,使奖励模型与ISR模型同步演进:奖励模型逐步提升复杂度,而ISR模型则从简单全局奖励开始快速收敛,再逐步过渡到复杂模型输出。这种由易到难的策略实现了稳定训练,并有效抑制奖励欺骗。实验验证了该方法在多种超分辨模型上的有效性,显著提升全局质量和局部真实感。

原文摘要 · Abstract (English)

Reinforcement Learning with Human Feedback (RLHF) has proven effective in image generation field guided by reward models to align human preferences. Motivated by this, adapting RLHF for Image Super-Resolution (ISR) tasks has shown promise in optimizing perceptual quality with Image Quality Assessment (IQA) model as reward models. However, the traditional IQA model usually output a single global score, which are exceptionally insensitive to local and fine-grained distortions. This insensitivity allows ISR models to produce perceptually undesirable artifacts that yield spurious high scores, misaligning optimization objectives with perceptual quality and results in reward hacking. To address this, we propose a Fine-grained Perceptual Reward Model (FinPercep-RM) based on an Encoder-Decoder architecture. While providing a global quality score, it also generates a Perceptual Degradation Map that spatially localizes and quantifies local defects. We specifically introduce the FGR-30k dataset to train this model, consisting of diverse and subtle distortions from real-world super-resolution models. Despite the success of the FinPercep-RM model, its complexity introduces significant challenges in generator policy learning, leading to training instability. To address this, we propose a Co-evolutionary Curriculum Learning (CCL) mechanism, where both the reward model and the ISR model undergo synchronized curricula. The reward model progressively increases in complexity, while the ISR model starts with a simpler global reward for rapid convergence, gradually transitioning to the more complex model outputs. This easy-to-hard strategy enables stable training while suppressing reward hacking. Experiments validates the effectiveness of our method across ISR models in both global quality and local realism on RLHF methods.

超分辨率强化学习奖励模型感知质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。