arXiv:2505.13152eess.IVcs.CV2025-05ECCV被引 3

用扩散模型提升图像视频压缩的感知质量,同时显著提高保真度。

Higher fidelity perceptual image and video compression with a latent conditioned residual denoising diffusion model

  • 先用解码器生成低失真初始图像,再用条件扩散模型预测残差优化感知质量。
  • 在标准测试中,PSNR最高提升2dB,感知质量与CDC相当。
  • 方法可轻松扩展至视频压缩,适合追求高感知质量的多媒体应用。

去噪扩散模型在图像生成任务中表现优异,常超越基于GAN的方法。近期,扩散模型被用于感知图像压缩(如CDC),但其主要缺点是虽能提供出色的感知质量,却在保真度上低于传统或学习型压缩方案。本文提出一种混合压缩方案,以优化感知质量,通过在CDC基础上引入解码网络,降低对保真度指标(如PSNR)的影响。首先由解码器生成初始图像(优化失真),随后潜变量条件扩散模型通过预测残差进一步精修重建,提升感知质量。在标准基准测试中,本方法实现高达+2dB的PSNR提升,同时保持与CDC相当的LPIPS和FID感知分数。该方法亦可轻松扩展至视频压缩,取得相似效果。

原文摘要 · Abstract (English)

Denoising diffusion models achieved impressive results on several image generation tasks often outperforming GAN based models. Recently, the generative capabilities of diffusion models have been employed for perceptual image compression, such as in CDC. A major drawback of these diffusion-based methods is that, while producing impressive perceptual quality images they are dropping in fidelity/increasing the distortion to the original uncompressed images when compared with other traditional or learned image compression schemes aiming for fidelity. In this paper, we propose a hybrid compression scheme optimized for perceptual quality, extending the approach of the CDC model with a decoder network in order to reduce the impact on distortion metrics such as PSNR. After using the decoder network to generate an initial image, optimized for distortion, the latent conditioned diffusion model refines the reconstruction for perceptual quality by predicting the residual. On standard benchmarks, we achieve up to +2dB PSNR fidelity improvements while maintaining comparable LPIPS and FID perceptual scores when compared with CDC. Additionally, the approach is easily extensible to video compression, where we achieve similar results.

图像压缩扩散模型感知质量视频压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。