arXiv:2501.10325cs.CV2025-01被引 2

提出首个用于双目图像修复的高频感知扩散模型,兼顾质量与效率。

DiffStereo: High-Frequency Aware Diffusion Model for Stereo Image Restoration

  • 在潜在空间学习高频率表征,保留纹理细节
  • 融合变压器网络,实现双目图像超分/去模糊/低光增强领先效果
  • 专为双目修复设计,适合图像重建与生成任务研究者

扩散模型(DMs)在图像修复中表现优异,但尚未应用于双目图像。其应用面临双重挑战:需同时重建两张图像导致计算开销大;现有潜在扩散模型常将高频细节视为冗余信息而丢弃,而这正是图像修复的关键。为此,我们首次提出高频率感知扩散模型 DiffStereo,用于双目图像修复。该方法首先学习高质量图像的潜在高频率表征(LHFR),并在该空间训练扩散模型以估计双目图像的LHFR,再将其融合至基于变压器的双目图像修复网络,提供来自对应高质量图像的有益高频信息。LHFR分辨率与输入图像一致,保留原始纹理,通道压缩则降低扩散模型计算负担。此外,我们设计了位置编码方案,在修复网络不同深度实现差异化引导。大量实验表明,通过结合生成式扩散模型与变压器结构,DiffStereo 在双目超分辨率、去模糊和低光增强任务上均优于当前最优方法,兼具更高的重建精度与更佳的感知质量。

原文摘要 · Abstract (English)

Diffusion models (DMs) have achieved promising performance in image restoration but haven't been explored for stereo images. The application of DM in stereo image restoration is confronted with a series of challenges. The need to reconstruct two images exacerbates DM's computational cost. Additionally, existing latent DMs usually focus on semantic information and remove high-frequency details as redundancy during latent compression, which is precisely what matters for image restoration. To address the above problems, we propose a high-frequency aware diffusion model, DiffStereo for stereo image restoration as the first attempt at DM in this domain. Specifically, DiffStereo first learns latent high-frequency representations (LHFR) of HQ images. DM is then trained in the learned space to estimate LHFR for stereo images, which are fused into a transformer-based stereo image restoration network providing beneficial high-frequency information of corresponding HQ images. The resolution of LHFR is kept the same as input images, which preserves the inherent texture from distortion. And the compression in channels alleviates the computational burden of DM. Furthermore, we devise a position encoding scheme when integrating the LHFR into the restoration network, enabling distinctive guidance in different depths of the restoration network. Comprehensive experiments verify that by combining generative DM and transformer, DiffStereo achieves both higher reconstruction accuracy and better perceptual quality on stereo super-resolution, deblurring, and low-light enhancement compared with state-of-the-art methods.

扩散模型双目图像图像修复高频感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。