用统一量化扩散模型实现渐进式图像压缩,单模型适配多码率。
Progressive Compression with Universally Quantized Diffusion Models
- 前向过程使用均匀噪声,通过负ELBO直接对应通用量化压缩成本。
- 在多个码率下达到竞争力的率失真与真实感性能。
- 适合需要高效、灵活压缩的神经编解码应用场景。
扩散概率模型在图像生成到逆问题求解等生成建模任务中取得主流成功。这类模型本质上是优化数据似然下界(ELBO)的深层层次隐变量模型。基于似然建模与压缩之间的基本联系,我们探索了扩散模型在渐进编码中的潜力,可生成逐级传输、解码后重建质量逐步提升的比特序列。不同于以往基于高斯或条件扩散模型的工作,本文提出一种前向过程采用均匀噪声的新式扩散模型,其负ELBO对应于使用通用量化时的端到端压缩代价。在图像压缩任务上获得初步优异结果,在多种码率下均表现出色,单个模型即可实现竞争性的率失真与率真实度表现,推动神经编解码向实际部署更进一步。
原文摘要 · Abstract (English)
Diffusion probabilistic models have achieved mainstream success in many generative modeling tasks, from image generation to inverse problem solving. A distinct feature of these models is that they correspond to deep hierarchical latent variable models optimizing a variational evidence lower bound (ELBO) on the data likelihood. Drawing on a basic connection between likelihood modeling and compression, we explore the potential of diffusion models for progressive coding, resulting in a sequence of bits that can be incrementally transmitted and decoded with progressively improving reconstruction quality. Unlike prior work based on Gaussian diffusion or conditional diffusion models, we propose a new form of diffusion model with uniform noise in the forward process, whose negative ELBO corresponds to the end-to-end compression cost using universal quantization. We obtain promising first results on image compression, achieving competitive rate-distortion and rate-realism results on a wide range of bit-rates with a single model, bringing neural codecs a step closer to practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。