让图像压缩自动匹配内容复杂度,实现超低码率下高保真重建。
CADC: Content Adaptive Diffusion-Based Generative Image Compression
- 根据图像局部复杂度动态调整量化精度,减少失真。
- 通过辅助解码器引导关键信息保留,提升重构质量。
- 用重建图自动生成语义提示,零码率代价实现内容感知指导。
基于扩散模型的生成式图像压缩在极低码率下展现出生成逼真图像的巨大潜力。其核心在于使整个压缩流程具备内容自适应能力,确保编码器表示与解码器生成先验能动态匹配输入图像的语义和结构特征。然而现有方法存在三大瓶颈:一是各向同性量化采用固定步长,无法适应图像内容的空间复杂度变化,导致与扩散模型依赖噪声的先验不匹配;二是信息集中瓶颈——由高维噪声潜在表示与解码器固定输入维度不匹配引发,阻碍模型自适应地在主通道中保留关键语义信息;三是现有文本条件策略或需高额文本码率,或依赖通用、非内容相关的提示,难以高效提供自适应语义引导。为此,本文提出一种内容自适应扩散生成图像编解码器,包含三项创新:1)不确定性引导的自适应量化方法,学习空间不确定性图以动态对齐量化失真与内容特征;2)辅助解码器引导的信息集中方法,利用轻量级辅助解码器强制在主潜在通道中实现内容感知的信息保留;3)无码率自适应文本条件方法,从辅助重建图像中提取内容感知的文本描述,实现语义引导而无需额外码率开销。
原文摘要 · Abstract (English)
Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process content-adaptive, ensuring that the encoder's representation and the decoder's generative prior are dynamically aligned with the semantic and structural characteristics of the input image. However, existing methods suffer from three critical limitations that prevent effective content adaptation. First, isotropic quantization applies a uniform quantization step, failing to adapt to the spatially varying complexity of image content and creating a misalignment with the diffusion model's noise-dependent prior. Second, the information concentration bottleneck -- arising from the dimensional mismatch between the high-dimensional noisy latent and the diffusion decoder's fixed input -- prevents the model from adaptively preserving essential semantic information in the primary channels. Third, existing textual conditioning strategies either need significant textual bitrate overhead or rely on generic, content-agnostic textual prompts, thereby failing to provide adaptive semantic guidance efficiently. To overcome these limitations, we propose a content-adaptive diffusion-based image codec with three technical innovations: 1) an Uncertainty-Guided Adaptive Quantization method that learns spatial uncertainty maps to adaptively align quantization distortion with content characteristics; 2) an Auxiliary Decoder-Guided Information Concentration method that uses a lightweight auxiliary decoder to enforce content-aware information preservation in the primary latent channels; and 3) a Bitrate-Free Adaptive Textual Conditioning method that derives content-aware textual descriptions from the auxiliary reconstructed image, enabling semantic guidance without bitrate cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。