arXiv:2510.18851cs.CVcs.AI2025-10NeurIPS被引 2

无需人工标注,用偏好优化提升真实图像超分辨率的视觉质量。

DP$^2$O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution

  • 直接优化感知偏好,利用生成模型输出的多样性构建奖励信号。
  • 在多个真实图像超分数据集上显著提升视觉质量,优于现有方法。
  • 适合关注图像生成质量与真实场景适配的研究者与工程师。

得益于预训练的文本到图像(T2I)扩散模型,真实世界图像超分辨率(Real-ISR)方法能够合成丰富且逼真的细节。然而,由于T2I模型固有的随机性,不同噪声输入常导致输出感知质量不一。这种随机性虽被视为局限,但也带来了更广的感知质量范围,可被用于提升性能。为此,我们提出直接感知偏好优化框架DP²O-SR,无需昂贵的人工标注即可对生成模型进行感知对齐。通过结合全参考和无参考图像质量评估模型构建混合奖励信号,该信号同时鼓励结构保真度与自然外观。为更好利用感知多样性,我们超越标准的最佳-最差选择,从同一模型输出中构建多个偏好对。分析表明,最优选择比例依赖于模型容量:小模型受益于更广覆盖,大模型则更适应更强对比监督。此外,我们提出分层偏好优化,基于组内奖励差距与组间多样性自适应加权训练样本,实现更高效稳定的训练。大量实验表明,无论采用扩散或流式T2I骨干网络,DP²O-SR均显著提升感知质量,并在真实世界基准上表现良好。

原文摘要 · Abstract (English)

Benefiting from pre-trained text-to-image (T2I) diffusion models, real-world image super-resolution (Real-ISR) methods can synthesize rich and realistic details. However, due to the inherent stochasticity of T2I models, different noise inputs often lead to outputs with varying perceptual quality. Although this randomness is sometimes seen as a limitation, it also introduces a wider perceptual quality range, which can be exploited to improve Real-ISR performance. To this end, we introduce Direct Perceptual Preference Optimization for Real-ISR (DP$^2$O-SR), a framework that aligns generative models with perceptual preferences without requiring costly human annotations. We construct a hybrid reward signal by combining full-reference and no-reference image quality assessment (IQA) models trained on large-scale human preference datasets. This reward encourages both structural fidelity and natural appearance. To better utilize perceptual diversity, we move beyond the standard best-vs-worst selection and construct multiple preference pairs from outputs of the same model. Our analysis reveals that the optimal selection ratio depends on model capacity: smaller models benefit from broader coverage, while larger models respond better to stronger contrast in supervision. Furthermore, we propose hierarchical preference optimization, which adaptively weights training pairs based on intra-group reward gaps and inter-group diversity, enabling more efficient and stable learning. Extensive experiments across both diffusion- and flow-based T2I backbones demonstrate that DP$^2$O-SR significantly improves perceptual quality and generalizes well to real-world benchmarks.

图像超分扩散模型偏好优化真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。