arXiv:2506.16572eess.IVcs.CV2025-06被引 2

单步扩散模型实现超低码率下高速高质图像压缩

Single-step Diffusion for Image Compression at Ultra-Low Bitrates

  • 用残差分解结构码与学习残差,分离图像几何与细节
  • 通过码率感知噪声调节,实现不同码率下的精准重建
  • 解码速度比以往扩散模型快50倍,适合实际应用

尽管图像压缩技术(如标准与学习编码器)取得显著进展,但在极低比特/像素下仍存在严重质量下降。现有基于扩散的模型虽在低码率下提升生成效果,但因需多步去噪导致感知质量有限且解码延迟过高。本文提出单步扩散图像压缩方法,在超低码率下实现高感知质量与快速解码。核心创新包括:(i) 向量量化残差训练(VQ-Residual),在潜在空间中分解结构基底码与学习残差,同时捕捉全局几何与高频细节;(ii) 码率感知噪声调制,动态调整去噪强度以匹配目标码率。大量实验表明,本方法在压缩性能上媲美当前最优水平,解码速度相较之前扩散方法提升约50倍,显著提升生成编码器的实用性。

原文摘要 · Abstract (English)

Although there have been significant advancements in image compression techniques, such as standard and learned codecs, these methods still suffer from severe quality degradation at extremely low bits per pixel. While recent diffusion-based models provided enhanced generative performance at low bitrates, they often yields limited perceptual quality and prohibitive decoding latency due to multiple denoising steps. In this paper, we propose the single-step diffusion model for image compression that delivers high perceptual quality and fast decoding at ultra-low bitrates. Our approach incorporates two key innovations: (i) Vector-Quantized Residual (VQ-Residual) training, which factorizes a structural base code and a learned residual in latent space, capturing both global geometry and high-frequency details; and (ii) rate-aware noise modulation, which tunes denoising strength to match the desired bitrate. Extensive experiments show that ours achieves comparable compression performance to state-of-the-art methods while improving decoding speed by about 50x compared to prior diffusion-based methods, greatly enhancing the practicality of generative codecs.

图像压缩扩散模型单步生成超低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。