arXiv:2410.22830eess.IVcs.CV2024-10被引 11

提出可连续缩放的遥感图像超分辨率模型,效率显著提升。

Latent Diffusion, Implicit Amplification: Efficient Continuous-Scale Super-Resolution for Remote Sensing Images

  • 两阶段潜空间扩散架构,分离图像差异先验与重建过程
  • 支持非整数倍缩放,推理速度接近传统方法
  • 专为遥感图像设计,视觉质量与指标双优

扩散模型在超分辨率(SR)任务中表现优异,但现有方法常混淆SR与通用图像生成的本质差异。通用生成从零构建图像,而SR仅需恢复低分辨率(LR)图像缺失的高频细节。这一误解不仅增加训练难度,也降低推理效率。此外,现有基于扩散的SR方法通常仅支持固定整数倍缩放,缺乏对非整数倍缩放的灵活性。为此,本文提出面向遥感图像连续尺度超分辨率的高效弹性模型E²DiffSR。该模型采用两阶段潜空间扩散框架:第一阶段训练自编码器,捕捉高分辨率(HR)与LR图像间的差异先验;编码器忽略原始LR内容以减轻负担,解码器引入连续尺度上采样模块,借助差异先验完成重建。第二阶段在潜空间学习条件扩散模型,预测真实差异先验编码。实验表明,E²DiffSR在客观指标和视觉质量上均优于当前最优方法,且将扩散型SR的推理时间降至与非扩散方法相当水平。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have significantly improved performance in super-resolution (SR) tasks. However, previous research often overlooks the fundamental differences between SR and general image generation. General image generation involves creating images from scratch, while SR focuses specifically on enhancing existing low-resolution (LR) images by adding typically missing high-frequency details. This oversight not only increases the training difficulty but also limits their inference efficiency. Furthermore, previous diffusion-based SR methods are typically trained and inferred at fixed integer scale factors, lacking flexibility to meet the needs of up-sampling with non-integer scale factors. To address these issues, this paper proposes an efficient and elastic diffusion-based SR model (E$^2$DiffSR), specially designed for continuous-scale SR in remote sensing imagery. E$^2$DiffSR employs a two-stage latent diffusion paradigm. During the first stage, an autoencoder is trained to capture the differential priors between high-resolution (HR) and LR images. The encoder intentionally ignores the existing LR content to alleviate the encoding burden, while the decoder introduces an SR branch equipped with a continuous scale upsampling module to accomplish the reconstruction under the guidance of the differential prior. In the second stage, a conditional diffusion model is learned within the latent space to predict the true differential prior encoding. Experimental results demonstrate that E$^2$DiffSR achieves superior objective metrics and visual quality compared to the state-of-the-art SR methods. Additionally, it reduces the inference time of diffusion-based SR methods to a level comparable to that of non-diffusion methods.

超分辨率扩散模型遥感图像连续缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。