让图像超分更符合人眼偏好,减少幻觉和伪影。
DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
- 用语义引导的直接偏好优化,对齐人眼感知偏好。
- 在真实图像超分中提升质量,减少局部异常影响。
- 适合需要高视觉保真度的图像生成场景。
扩散模型在真实世界图像超分辨率(Real-ISR)中取得进展,但现有方法缺乏人类反馈整合,易导致与人类偏好偏离,产生伪影、幻觉甚至有害内容。为此,我们首次将人类偏好对齐引入 Real-ISR,借鉴大语言模型与文生图任务中的成功经验。具体地,采用直接偏好优化(DPO)实现对齐,但因 Real-ISR 的像素级重建目标与 DPO 的图像级偏好存在冲突,导致对局部异常敏感,影响生成质量。为解决此矛盾,提出直接语义偏好优化(DSPO),通过两项策略实现实例级偏好对齐:(a) 语义实例对齐策略,保证细粒度感知一致性;(b) 用户描述反馈策略,基于实例级文本反馈抑制幻觉。作为即插即用方案,DSPO 在单步与多步超分框架中均表现优异。
原文摘要 · Abstract (English)
Recent advances in diffusion models have improved Real-World Image Super-Resolution (Real-ISR), but existing methods lack human feedback integration, risking misalignment with human preference and may leading to artifacts, hallucinations and harmful content generation. To this end, we are the first to introduce human preference alignment into Real-ISR, a technique that has been successfully applied in Large Language Models and Text-to-Image tasks to effectively enhance the alignment of generated outputs with human preferences. Specifically, we introduce Direct Preference Optimization (DPO) into Real-ISR to achieve alignment, where DPO serves as a general alignment technique that directly learns from the human preference dataset. Nevertheless, unlike high-level tasks, the pixel-level reconstruction objectives of Real-ISR are difficult to reconcile with the image-level preferences of DPO, which can lead to the DPO being overly sensitive to local anomalies, leading to reduced generation quality. To resolve this dichotomy, we propose Direct Semantic Preference Optimization (DSPO) to align instance-level human preferences by incorporating semantic guidance, which is through two strategies: (a) semantic instance alignment strategy, implementing instance-level alignment to ensure fine-grained perceptual consistency, and (b) user description feedback strategy, mitigating hallucinations through semantic textual feedback on instance-level images. As a plug-and-play solution, DSPO proves highly effective in both one-step and multi-step SR frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。