GuideSR用双分支结构提升单步超分辨率图像保真度。
GuideSR: Rethinking Guidance for One-Step High-Fidelity Diffusion-Based Super-Resolution
- 双分支设计:保留原始输入结构,用扩散模型提升感知质量。
- 在真实数据集上提升1.39dB PSNR,超越现有方法。
- 适合追求高保真、低计算成本的图像恢复应用。
本文提出GuideSR,一种新型单步扩散超分辨率模型,专为提升图像保真度而设计。现有扩散超分辨率方法通常通过在降采样输入的VAE表示上添加额外条件来适配生成模型,但常损害结构保真度。GuideSR通过双分支架构解决此问题:(1)引导分支保留原始分辨率退化输入中的高保真结构;(2)扩散分支利用预训练潜空间扩散模型增强感知质量。不同于传统条件机制,引导分支采用针对图像修复任务定制的结构,结合全分辨率块(FRBs)与通道注意力,以及带引导注意力的图像引导网络(IGN)。通过将详细结构信息直接嵌入恢复流程,GuideSR生成更锐利且视觉一致的结果。大量基准测试表明,GuideSR在保持单步方法低计算开销的同时达到最先进性能,在具有挑战性的真实世界数据集上实现最高1.39dB的PSNR提升。其在多种参考基评估指标(包括PSNR、SSIM、LPIPS、DISTS和FID)上持续优于现有方法,为真实世界图像恢复提供了实际进展。
原文摘要 · Abstract (English)
In this paper, we propose GuideSR, a novel single-step diffusion-based image super-resolution (SR) model specifically designed to enhance image fidelity. Existing diffusion-based SR approaches typically adapt pre-trained generative models to image restoration tasks by adding extra conditioning on a VAE-downsampled representation of the degraded input, which often compromises structural fidelity. GuideSR addresses this limitation by introducing a dual-branch architecture comprising: (1) a Guidance Branch that preserves high-fidelity structures from the original-resolution degraded input, and (2) a Diffusion Branch, which a pre-trained latent diffusion model to enhance perceptual quality. Unlike conventional conditioning mechanisms, our Guidance Branch features a tailored structure for image restoration tasks, combining Full Resolution Blocks (FRBs) with channel attention and an Image Guidance Network (IGN) with guided attention. By embedding detailed structural information directly into the restoration pipeline, GuideSR produces sharper and more visually consistent results. Extensive experiments on benchmark datasets demonstrate that GuideSR achieves state-of-the-art performance while maintaining the low computational cost of single-step approaches, with up to 1.39dB PSNR gain on challenging real-world datasets. Our approach consistently outperforms existing methods across various reference-based metrics including PSNR, SSIM, LPIPS, DISTS and FID, further representing a practical advancement for real-world image restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。