arXiv:2505.08281cs.CVeess.IV2025-05ICML被引 18

用语义残差提升超低码率图像压缩,重建质量更优。

Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion

  • 引入语义残差编码,捕捉原图与压缩表示的语义差异
  • 在LPIPS和FID上实现-80.7%和-66.3%的比特率节省
  • 专为压缩优化的扩散模型,提升压缩与重建协同效果

现有基于多模态大模型的图像压缩框架常依赖语义检索、潜在空间压缩与生成模型的碎片化集成,导致重建保真度与编码效率均不理想。为此,我们提出一种残差引导的超低码率图像压缩方法ResULIC,将残差信号融入语义检索与基于扩散的生成过程。具体而言,引入语义残差编码(SRC)以捕捉原始图像与其压缩潜在表示之间的语义差异,并通过感知保真度优化器提升重建质量。此外,提出压缩感知扩散模型(CDM),实现码率与扩散时间步的最优对齐,增强压缩-重建协同性。大量实验表明,ResULIC在客观与主观性能上均优于现有扩散基方法,在LPIPS和FID指标上分别实现-80.7%和-66.3%的BD-rate降低。项目主页见 https://njuvision.github.io/ResULIC/。

原文摘要 · Abstract (English)

Existing multimodal large model-based image compression frameworks often rely on a fragmented integration of semantic retrieval, latent compression, and generative models, resulting in suboptimal performance in both reconstruction fidelity and coding efficiency. To address these challenges, we propose a residual-guided ultra lowrate image compression named ResULIC, which incorporates residual signals into both semantic retrieval and the diffusion-based generation process. Specifically, we introduce Semantic Residual Coding (SRC) to capture the semantic disparity between the original image and its compressed latent representation. A perceptual fidelity optimizer is further applied for superior reconstruction quality. Additionally, we present the Compression-aware Diffusion Model (CDM), which establishes an optimal alignment between bitrates and diffusion time steps, improving compression-reconstruction synergy. Extensive experiments demonstrate the effectiveness of ResULIC, achieving superior objective and subjective performance compared to state-of-the-art diffusion-based methods with - 80.7%, -66.3% BD-rate saving in terms of LPIPS and FID. Project page is available at https: //njuvision.github.io/ResULIC/.

图像压缩扩散模型低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。