用双向奖励引导扩散模型,让真实图像超分辨率更清晰自然。
Bird-SR: Bidirectional Reward-Guided Diffusion for Real-World Image Super-Resolution
- 通过奖励反馈学习,联合使用合成与真实图像优化
- 在真实图像上实现更高感知质量且结构一致
- 适合追求真实感图像修复的视觉任务使用者
基于多模态文本到图像先验的扩散超分辨率在生成细节方面表现优异;然而,仅在合成低分辨率(LR)与高分辨率(HR)图像对上训练的模型,在处理真实世界LR图像时因分布差异导致性能下降。本文提出Bird-SR,一种双向奖励引导的扩散框架,将超分辨率建模为轨迹级偏好优化,结合合成LR-HR对与真实世界LR图像进行联合训练。为保障结构保真度,模型在早期扩散步骤中直接优化于合成数据,从而减小真实输入在结构层面的分布差距。后期引入感知质量引导奖励,同时对合成结果采用相对优势空间的奖励设计,并对真实图像施加语义对齐约束以防止奖励欺骗。为平衡结构与感知学习,提出动态保真-感知权重策略:前期侧重结构保留,后期逐步转向感知优化。大量实验证明,Bird-SR在真实世界超分辨率基准上持续优于现有方法,在感知质量与结构一致性之间取得良好平衡。
原文摘要 · Abstract (English)
Powered by multimodal text-to-image priors, diffusion-based super-resolution excels at synthesizing intricate details; however, models trained on synthetic low-resolution (LR) and high-resolution (HR) image pairs often degrade when applied to real-world LR images due to significant distribution shifts. We propose Bird-SR, a bidirectional reward-guided diffusion framework that formulates super-resolution as trajectory-level preference optimization via reward feedback learning (ReFL), jointly leveraging synthetic LR-HR pairs and real-world LR images. For structural fidelity easily affected in ReFL, the model is directly optimized on synthetic pairs at early diffusion steps, which also facilitates structure preservation for real-world inputs under smaller distribution gap in structure levels. For perceptual enhancement, quality-guided rewards are applied to both synthetic and real LR images at the later trajectory phase. To mitigate reward hacking, the rewards for synthetic results are formulated in a relative advantage space bounded by their ground-truth counterparts, while real-world optimization is regularized via a semantic alignment constraint. Furthermore, to balance structural and perceptual learning, we introduce a dynamic fidelity-perception weighting strategy that emphasizes structure preservation at early stages and progressively shifts focus toward perceptual optimization at later diffusion steps. Extensive experiments on real-world SR benchmarks demonstrate that Bird-SR consistently outperforms state-of-the-art methods in perceptual quality while preserving structural consistency, validating its effectiveness for real-world super-resolution. Our code can be obtained at https://github.com/fanzh03/Bird-SR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。