arXiv:2506.02605cs.CV2025-06被引 6

一拍即合:用单步扩散模型实现真实图像超分辨率,画质更逼真。

One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation

  • 用CLIP提取语义监督,提升生成图像与真实图像的一致性。
  • 新增高频感知损失,恢复关键细节,使画质更清晰。
  • 适合追求高真实感、低延迟的图像修复与增强场景。

基于扩散模型的图像超分辨率虽表现优异,但通常需数十甚至上百次采样步骤。现有方法通过知识蒸馏加速推理,却在语义对齐和感知质量上存在不足,尤其体现在CLIPIQA得分偏低。为此,我们提出一种面向超分辨率的视觉感知蒸馏框架VPD-SR,实现仅一步采样的高效生成。该方法包含显式语义监督(ESS)和高频感知(HFP)损失:前者利用CLIP模型提取语义指导,提升一致性;后者引导学生模型恢复退化图像中缺失的高频细节,显著改善感知质量。此外,通过对抗训练进一步增强生成内容的真实性。在合成与真实世界数据集上的实验表明,仅需一步采样,VPD-SR性能优于此前最先进方法及教师模型。

原文摘要 · Abstract (English)

Diffusion-based models have been widely used in various visual generation tasks, showing promising results in image super-resolution (SR), while typically being limited by dozens or even hundreds of sampling steps. Although existing methods aim to accelerate the inference speed of multi-step diffusion-based SR methods through knowledge distillation, their generated images exhibit insufficient semantic alignment with real images, resulting in suboptimal perceptual quality reconstruction, specifically reflected in the CLIPIQA score. These methods still have many challenges in perceptual quality and semantic fidelity. Based on the challenges, we propose VPD-SR, a novel visual perception diffusion distillation framework specifically designed for SR, aiming to construct an effective and efficient one-step SR model. Specifically, VPD-SR consists of two components: Explicit Semantic-aware Supervision (ESS) and High-Frequency Perception (HFP) loss. Firstly, the ESS leverages the powerful visual perceptual understanding capabilities of the CLIP model to extract explicit semantic supervision, thereby enhancing semantic consistency. Then, Considering that high-frequency information contributes to the visual perception quality of images, in addition to the vanilla distillation loss, the HFP loss guides the student model to restore the missing high-frequency details in degraded images that are critical for enhancing perceptual quality. Lastly, we expand VPD-SR in adversarial training manner to further enhance the authenticity of the generated content. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed VPD-SR achieves superior performance compared to both previous state-of-the-art methods and the teacher model with just one-step sampling.

图像超分扩散模型感知质量单步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。