arXiv:2411.18824cs.CV2024-11CVPR被引 37

用扩散模型提升图像超分辨率,让恢复图像既真实又忠实原图结构。

FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolution

  • 利用扩散模型先验,动态提取输入图像中的有用特征以恢复结构。
  • 在多个数据集上优于当前最佳方法,尤其在保持图像细节和结构一致性上表现突出。
  • 适合需要高保真度超分辨率的视觉任务,如医学影像、历史图像修复。

忠实图像超分辨率(SR)不仅要求恢复出视觉上逼真的图像,还需保证重建结果与输入图像在结构和语义上保持一致。为此,本文提出一种简单而有效的方法 FaithDiff,旨在充分挖掘潜在扩散模型(LDMs)在忠实重建中的潜力。与现有基于扩散的 SR 方法冻结预训练扩散模型不同,FaithDiff 解放扩散先验,主动识别并恢复关键结构信息。由于退化输入特征与扩散模型中的噪声潜在表示之间存在显著差异,我们设计了一个有效的对齐模块,用于从退化输入中提取有用特征,并使其与扩散过程相匹配。考虑到编码器与扩散模型在 LDM 中的协同作用,我们在统一优化框架中联合微调二者,使编码器能提取更契合扩散过程的特征。大量实验表明,FaithDiff 在多个基准数据集上超越当前最优方法,生成高质量且高度忠实的超分辨率图像。

原文摘要 · Abstract (English)

Faithful image super-resolution (SR) not only needs to recover images that appear realistic, similar to image generation tasks, but also requires that the restored images maintain fidelity and structural consistency with the input. To this end, we propose a simple and effective method, named FaithDiff, to fully harness the impressive power of latent diffusion models (LDMs) for faithful image SR. In contrast to existing diffusion-based SR methods that freeze the diffusion model pre-trained on high-quality images, we propose to unleash the diffusion prior to identify useful information and recover faithful structures. As there exists a significant gap between the features of degraded inputs and the noisy latent from the diffusion model, we then develop an effective alignment module to explore useful features from degraded inputs to align well with the diffusion process. Considering the indispensable roles and interplay of the encoder and diffusion model in LDMs, we jointly fine-tune them in a unified optimization framework, facilitating the encoder to extract useful features that coincide with diffusion process. Extensive experimental results demonstrate that FaithDiff outperforms state-of-the-art methods, providing high-quality and faithful SR results.

图像超分扩散模型结构保真生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。