用四元数小波增强扩散模型,提升图像超分辨率细节质量。
Quaternion Wavelet-Conditioned Diffusion Models for Image Super-Resolution
- 通过四元数小波嵌入动态调节去噪过程,强化条件输入。
- 在多个数据集上超越现有方法,尤其在高倍率放大时保持纹理真实。
- 适合需要高保真图像重建的医学影像与遥感分析场景。
图像超分辨率是计算机视觉中的基础问题,广泛应用于医学成像和卫星分析等领域。从低分辨率输入重建高质量高分辨率图像对下游任务如目标检测和分割至关重要。尽管深度学习已显著推进该领域,但在高倍率放大下实现精细细节和真实纹理仍具挑战。近期基于扩散模型的方法虽有进展,但常难以兼顾感知质量与结构保真度。本文提出ResQu框架,将四元数小波预处理与潜在扩散模型结合,引入新的四元数小波-时间感知编码器。不同于简单在扩散模型中加入小波变换,本方法通过动态整合四元数小波嵌入,优化不同去噪阶段的条件信号。同时利用Stable Diffusion等基础模型的生成先验。在多个领域特定数据集上的实验证明,该方法在感知质量与标准评估指标上均优于现有方法。代码已开源。
原文摘要 · Abstract (English)
Image Super-Resolution is a fundamental problem in computer vision with broad applications spacing from medical imaging to satellite analysis. The ability to reconstruct high-resolution images from low-resolution inputs is crucial for enhancing downstream tasks such as object detection and segmentation. While deep learning has significantly advanced SR, achieving high-quality reconstructions with fine-grained details and realistic textures remains challenging, particularly at high upscaling factors. Recent approaches leveraging diffusion models have demonstrated promising results, yet they often struggle to balance perceptual quality with structural fidelity. In this work, we introduce ResQu a novel SR framework that integrates a quaternion wavelet preprocessing framework with latent diffusion models, incorporating a new quaternion wavelet- and time-aware encoder. Unlike prior methods that simply apply wavelet transforms within diffusion models, our approach enhances the conditioning process by exploiting quaternion wavelet embeddings, which are dynamically integrated at different stages of denoising. Furthermore, we also leverage the generative priors of foundation models such as Stable Diffusion. Extensive experiments on domain-specific datasets demonstrate that our method achieves outstanding SR results, outperforming in many cases existing approaches in perceptual quality and standard evaluation metrics. The code is available at https://www.github.com/Fascetta/ResQu
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。