用扩散模型实现超低码率图像压缩,生成效果逼真。
Advances in Diffusion-Based Generative Compression
- 编码时生成嵌入,解码时用扩散模型逐步优化重建
- 可在极低码率下生成接近真实分布的图像
- 揭示了随机性共享与反问题求解的深层联系
受其强大的图像生成能力推动,扩散模型及其相关方法在视觉媒体应用中广受欢迎。特别是,扩散模型已催生新型有损数据压缩方法,可在极低比特率下生成逼真的重构图像。本文对近期基于扩散模型的生成式有损压缩方法进行了统一综述,重点关注图像压缩。这些方法通常将源数据编码为嵌入表示,并在解码过程中利用扩散模型迭代精炼该嵌入,使最终重建结果近似于真实数据分布。嵌入可采用多种形式,通常通过辅助熵模型传输;近期方法还探索使用扩散模型自身进行信息传输,通过信道模拟实现。我们从率-失真-感知理论视角回顾代表性方法,强调共同随机性的作用及与反问题的关联,并指出当前开放挑战。
原文摘要 · Abstract (English)
Popularized by their strong image generation performance, diffusion and related methods for generative modeling have found widespread success in visual media applications. In particular, diffusion methods have enabled new approaches to data compression, where realistic reconstructions can be generated at extremely low bit-rates. This article provides a unifying review of recent diffusion-based methods for generative lossy compression, with a focus on image compression. These methods generally encode the source into an embedding and employ a diffusion model to iteratively refine it in the decoding procedure, such that the final reconstruction approximately follows the ground truth data distribution. The embedding can take various forms and is typically transmitted via an auxiliary entropy model, and recent methods also explore the use of diffusion models themselves for information transmission via channel simulation. We review representative approaches through the lens of rate-distortion-perception theory, highlighting the role of common randomness and connections to inverse problems, and identify open challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。