用新量化方法让扩散模型在极低码率下生成更真实图像
Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compression
- 设计统一量化方案与定制化量化调度,修复三类分布偏差
- 在极低码率下实现最佳保真度与真实感平衡,优于现有方法
- 适合需要高保真图像压缩的低带宽场景应用
生成式神经图像压缩可在极低码率下表示数据,通过客户端合成细节并持续生成高度逼真的图像。利用量化误差与加性噪声的相似性,可基于潜在扩散模型构建生成式压缩编码器,对量化引入的伪影进行“去噪”。然而,我们识别出此前遵循此范式的三种关键差距(噪声水平、噪声类型与离散化差距),导致量化数据偏离扩散模型所知的数据分布。本文提出一种具有理论基础的新量化前向扩散过程,同时解决上述三类问题。通过精心设计的通用量化方案与均匀噪声训练的扩散模型,实现一致且精细的重建效果,即使在极低码率下也显著优于先前工作,达成最优的率失真-真实感性能表现。
原文摘要 · Abstract (English)
Generative neural image compression supports data representation at extremely low bitrate, synthesizing details at the client and consistently producing highly realistic images. By leveraging the similarities between quantization error and additive noise, diffusion-based generative image compression codecs can be built using a latent diffusion model to "denoise" the artifacts introduced by quantization. However, we identify three critical gaps in previous approaches following this paradigm (namely, the noise level, noise type, and discretization gaps) that result in the quantized data falling out of the data distribution known by the diffusion model. In this work, we propose a novel quantization-based forward diffusion process with theoretical foundations that tackles all three aforementioned gaps. We achieve this through universal quantization with a carefully tailored quantization schedule and a diffusion model trained with uniform noise. Compared to previous work, our proposal produces consistently realistic and detailed reconstructions, even at very low bitrates. In such a regime, we achieve the best rate-distortion-realism performance, outperforming previous related works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。