arXiv:2504.08591cs.CV2025-04被引 5

ZipIR通过压缩潜空间实现超高清图像修复,兼顾速度与质量。

ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration

  • 用潜空间金字塔VAE将图像压缩32倍,减少计算量
  • 在2K分辨率下修复严重退化图像,速度与质量均领先
  • 适合需要高效高保真图像恢复的研究与应用

生成模型的进展显著提升了图像修复能力,尤其是扩散模型在语义细节和局部保真度方面表现突出。然而,超高清图像修复面临质量和效率的权衡,主要源于长程注意力机制的计算开销。为此,我们提出ZipIR,一种提升效率、可扩展性及长程建模能力的新框架。ZipIR采用高度压缩的潜表示,将图像压缩32倍,有效减少空间令牌数量,并支持高容量模型如Diffusion Transformer(DiT)。为此,我们设计了潜空间金字塔VAE(LP-VAE),将潜空间分层为子带,以简化扩散训练。在高达2K分辨率的完整图像上训练,ZipIR超越现有基于扩散的方法,在严重退化输入的高分辨率图像修复中展现出前所未有的速度与质量。

原文摘要 · Abstract (English)

Recent progress in generative models has significantly improved image restoration capabilities, particularly through powerful diffusion models that offer remarkable recovery of semantic details and local fidelity. However, deploying these models at ultra-high resolutions faces a critical trade-off between quality and efficiency due to the computational demands of long-range attention mechanisms. To address this, we introduce ZipIR, a novel framework that enhances efficiency, scalability, and long-range modeling for high-res image restoration. ZipIR employs a highly compressed latent representation that compresses image 32x, effectively reducing the number of spatial tokens, and enabling the use of high-capacity models like the Diffusion Transformer (DiT). Toward this goal, we propose a Latent Pyramid VAE (LP-VAE) design that structures the latent space into sub-bands to ease diffusion training. Trained on full images up to 2K resolution, ZipIR surpasses existing diffusion-based methods, offering unmatched speed and quality in restoring high-resolution images from severely degraded inputs.

图像修复扩散模型潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。