arXiv:2505.16177eess.IVcs.CV2025-05被引 31

用生成模型隐空间编码,实现超低码率下高清图像视频压缩

Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression

  • 在生成模型的隐空间进行变换编码,更贴合人眼感知
  • 图像压缩低于0.04 bpp,FID表现优于现有最优模型
  • 适合追求极致压缩比的视觉应用,如移动端传输

现有图像和视频压缩方法通常在像素空间进行变换编码以减少冗余,但因像素空间失真与人类感知不匹配,难以在超低码率下同时实现高真实感和高保真。为此,本文提出生成隐空间编码(GLC)模型,包括GLC-image和GLC-video。其变换编码在生成式向量量化变分自编码器(VQ-VAE)的隐空间中进行,该空间具有更强稀疏性、更丰富的语义信息,且更契合人类感知,从而在高真实感与高保真度方面表现优异。为进一步提升性能,我们在GLC-image中引入空间类别先验模块,在GLC-video中引入时空类别先验模块;并提出基于码本预测的损失函数以增强语义一致性。实验表明,该方案在超低码率下仍保持高视觉质量:图像压缩在CLIC 2020测试集上达到小于0.04 bpp的比特率,与SOTA模型MS-ILLM相比在相同FID下节省45%码率;视频压缩在DISTS指标上相较PLVC节省65.3%码率。

原文摘要 · Abstract (English)

Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose \textbf{G}enerative \textbf{L}atent \textbf{C}oding (\textbf{GLC}) models for image and video compression, termed GLC-image and GLC-Video. The transform coding of GLC is conducted in the latent space of a generative vector-quantized variational auto-encoder (VQ-VAE). Compared to the pixel-space, such a latent space offers greater sparsity, richer semantics and better alignment with human perception, and show its advantages in achieving high-realism and high-fidelity compression. To further enhance performance, we improve the hyper prior by introducing a spatial categorical hyper module in GLC-image and a spatio-temporal categorical hyper module in GLC-video. Additionally, the code-prediction-based loss function is proposed to enhance the semantic consistency. Experiments demonstrate that our scheme shows high visual quality at ultra-low bitrate for both image and video compression. For image compression, GLC-image achieves an impressive bitrate of less than $0.04$ bpp, achieving the same FID as previous SOTA model MS-ILLM while using $45\%$ fewer bitrate on the CLIC 2020 test set. For video compression, GLC-video achieves 65.3\% bitrate saving over PLVC in terms of DISTS.

图像压缩视频压缩生成模型隐空间编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。