arXiv:2505.12935cs.CV2025-05被引 1

用可逆网络模拟图像退化,实现无固定模型的高效修复。

LatentINDIGO: An INN-Guided Latent Diffusion Algorithm for Image Restoration

  • 设计波浪变换启发的可逆神经网络,模拟退化并恢复细节。
  • 在合成与真实低质图像上达到当前最优效果,支持任意输出尺寸。
  • 全潜空间运行,降低计算开销,适合未知退化场景使用。

由于能有效建模自然图像分布,潜在扩散模型(LDMs)在图像修复(IR)任务中受到越来越多关注。尽管取得进展,仍存在关键挑战:首先,许多方法依赖预定义退化算子,难以应对复杂或未知退化;其次,多数方法在潜空间中难以稳定引导;最后,多数方法需在每采样迭代中将潜表示转回像素域进行引导,显著增加计算与内存开销。为此,我们提出一种受小波启发的可逆神经网络(INN),通过前向变换模拟退化,反向变换重建丢失细节。进一步将该设计融入潜扩散流程,提出两种方法:LatentINDIGO-PixelINN(在像素域操作)和LatentINDIGO-LatentINN(全程在潜空间运行,降低复杂度)。两者交替更新中间潜变量,并优化INN前向模型以处理未知退化。此外,正则化步骤保持潜变量贴近自然图像流形。实验表明,本算法在合成与真实低质图像上均达当前最优性能,且可灵活适配任意输出尺寸。

原文摘要 · Abstract (English)

There is a growing interest in the use of latent diffusion models (LDMs) for image restoration (IR) tasks due to their ability to model effectively the distribution of natural images. While significant progress has been made, there are still key challenges that need to be addressed. First, many approaches depend on a predefined degradation operator, making them ill-suited for complex or unknown degradations that deviate from standard analytical models. Second, many methods struggle to provide a stable guidance in the latent space and finally most methods convert latent representations back to the pixel domain for guidance at every sampling iteration, which significantly increases computational and memory overhead. To overcome these limitations, we introduce a wavelet-inspired invertible neural network (INN) that simulates degradations through a forward transform and reconstructs lost details via the inverse transform. We further integrate this design into a latent diffusion pipeline through two proposed approaches: LatentINDIGO-PixelINN, which operates in the pixel domain, and LatentINDIGO-LatentINN, which stays fully in the latent space to reduce complexity. Both approaches alternate between updating intermediate latent variables under the guidance of our INN and refining the INN forward model to handle unknown degradations. In addition, a regularization step preserves the proximity of latent variables to the natural image manifold. Experiments demonstrate that our algorithm achieves state-of-the-art performance on synthetic and real-world low-quality images, and can be readily adapted to arbitrary output sizes.

图像修复扩散模型可逆网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。