arXiv:2604.01122eess.IV2026-04

根据视觉重要性自适应分配比特,提升图像压缩感知质量。

Region-Adaptive Generative Compression with Spatially Varying Diffusion Models

  • 基于空间可变扩散模型,按重要区域动态调整去噪强度。
  • 引入重要性图作为隐空间先验,显著改善率失真性能。
  • 适合关注视觉焦点与细节保真的图像压缩应用。

生成式图像编码器旨在优化感知质量,生成逼真且细节丰富的重建图像。然而,它们常忽略人类视觉的一个关键特性:我们倾向于关注视觉场景中的特定部分(如显著物体),而对其他区域关注度较低。理想的感知编码器应能利用这一特性,将更多表示容量分配给感知重要的区域。为此,我们提出一种支持图像内非均匀比特分配的区域自适应扩散图像编码器。设计了一种新型空间可变扩散模型,能够根据任意重要性图对每个像素进行不同程度的去噪。进一步发现,这些重要性图可作为潜在表示的有效先验,并将其整合到熵模型中,提升了率失真性能。基于上述贡献,我们的空间自适应扩散编码器在全图及区域感兴趣(ROI)掩码下的感知质量上均优于当前最先进的可控制区域基线方法。

原文摘要 · Abstract (English)

Generative image codecs aim to optimize perceptual quality, producing realistic and detailed reconstructions. However, they often overlook a key property of human vision: our tendency to focus on particular aspects of a visual scene (e.g., salient objects) while giving less importance to other regions. An ideal perceptual codec should be able to exploit this property by allocating more representational capacity to perceptually important areas. To this end, we propose a region-adaptive diffusion-based image codec that supports non-uniform bit allocation within an image. We design a novel spatially varying diffusion model capable of denoising varying amounts of noise per pixel according to arbitrary importance maps. We further identify that these maps can serve as effective priors on the latent representation, and integrate them into our entropy model, improving rate-distortion performance. Built on these contributions, our spatially-adaptive diffusion-based codec outperforms state-of-the-art ROI-controllable baselines in both full-image and ROI-masked perceptual quality.

图像压缩扩散模型感知质量自适应编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。