解决真实图像超分中生成幻觉问题,提升细节真实度。
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
- 设计多奖励强化学习框架,基于低分辨率输入评估生成质量
- 在多个真实场景下提升感知质量,同时保持与原始图像的一致性
- 适用于需要高真实感图像生成的视觉任务研究者
生成式真实世界图像超分辨率(Real-ISR)可从严重退化的低分辨率(LR)输入中合成逼真细节,但其随机采样易导致关键失败:输出虽清晰却违背LR证据,出现语义或结构幻觉。偏好强化学习(RL)天然适用,因每个LR输入可生成一组候选重建结果。然而,有效对齐面临三大耦合挑战:(i) 缺乏对退化鲁棒且对局部幻觉敏感的LR参考一致性信号;(ii) 滚动组优化瓶颈——在归一化前对异质奖励标量处理会压缩各目标对比度,削弱DiffusionNFT类奖励加权更新;(iii) 真实退化覆盖有限,限制滚动多样性与偏好信号质量。本文提出LucidNFT,一种用于流匹配型Real-ISR的多奖励RL框架。引入LucidConsistency,一种退化不变、幻觉敏感的LR参考评估器,通过内容一致退化池和原始内补硬负样本训练;采用解耦奖励归一化策略,在融合前保留每组LR条件下的目标间对比度;并构建LucidLR,一个大规模真实世界退化图像集合,用于鲁棒的RL微调。大量实验表明,LucidNFT在强流基Real-ISR基线上的感知质量显著提升,同时在多样真实场景中普遍保持与LR的参考一致性。
原文摘要 · Abstract (English)
Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs, yet its stochastic sampling makes a critical failure mode hard to avoid: outputs may look sharp but be unfaithful to the LR evidence, exhibiting semantic or structural hallucinations. Preference-based reinforcement learning (RL) is a natural fit because each LR input yields a rollout group of candidate restorations. However, effective alignment in Real-ISR is hindered by three coupled challenges: (i) the lack of an LR-referenced faithfulness signal that is robust to degradation yet sensitive to localized hallucinations, (ii) a rollout-group optimization bottleneck where scalarizing heterogeneous rewards before normalization compresses objective-wise contrasts and weakens DiffusionNFT-style reward-weighted updates, and (iii) limited coverage of real degradations, which restricts rollout diversity and preference signal quality. We propose LucidNFT, a multi-reward RL framework for flow-matching Real-ISR. LucidNFT introduces LucidConsistency, a degradation-invariant and hallucination-sensitive LR-referenced evaluator trained with content-consistent degradation pools and original-inpainted hard negatives; a decoupled reward normalization strategy that preserves objective-wise contrasts within each LR-conditioned rollout group before fusion; and LucidLR, a large-scale collection of real-world degraded images for robust RL fine-tuning. Extensive experiments show that LucidNFT improves perceptual quality on strong flow-based Real-ISR baselines while generally maintaining LR-referenced consistency across diverse real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。