用不确定性引导扩散过程,提升图像超分辨率的保真与感知质量平衡。
Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

- 基于潜在特征的重建不确定性,动态引导高频细节修复区域。
- 在保持整体保真度的同时,显著改善感知质量,优于当前最优扩散模型。
- 适合追求高质量图像生成的视觉任务研究者使用。
单图超分辨率面临感知与失真之间的权衡难题。尽管基于扩散的超分辨率方法在生成逼真图像方面表现优异,但高保真度仍是主要瓶颈。近期进展虽提升了保真度,但常因过度依赖高保真图像而牺牲感知质量。为此,本文提出UGDiff,一种新型扩散引导范式,通过估计高保真图像对应潜在特征的重建不确定性,指导扩散过程仅在高不确定性区域恢复高频细节,其余区域保持原有保真度。此外,该方法结合采样器每一步的后验方差,自适应识别高不确定性区域,降低后期采样对高保真图像的依赖,从而实现更优的感知-失真平衡。大量实验表明,本方法在多个基准上均优于现有先进扩散模型。
原文摘要 · Abstract (English)
The perception-distortion trade-off poses a fundamental challenge in single-image super-resolution (SR). Although diffusion-based SR methods excel at generating perceptually realistic images, achieving high fidelity remains a key limitation. Recent advances in diffusion-based SR have shown promise in improving fidelity, but these methods often compromise perceptual quality due to their high reliance on a high-fidelity image. To address this, we introduce UGDiff, a novel diffusion guidance paradigm designed to further improve the perception-distortion balance. In particular, we first estimate the reconstruction uncertainty of the latent features corresponding to a high-fidelity image. This uncertainty is then used to guide the diffusion process to selectively restore high-frequency details in high-uncertainty regions, while preserving fidelity elsewhere. Furthermore, our guidance method adaptively identifies the high-uncertainty regions by considering not only the estimated uncertainty but also the posterior variance of the diffusion sampler at each timestep. This relaxes the reliance on the high-fidelity image in the later stages of sampling, thereby achieving a better perception-distortion balance. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art diffusion-based SR methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。