arXiv:2503.11321cs.CVeess.IV2025-03被引 8

用扩散模型提升图像压缩的纹理真实感,兼顾画质与带宽效率。

Leveraging Diffusion Knowledge for Generative Image Compression with Fractal Frequency-Aware Band Learning

  • 设计分形频域感知网络,捕捉自然图像的方向性频率特征。
  • 在编码器融入扩散知识,解码时迭代恢复丢失纹理细节。
  • 结合频域与内容感知正则化,实现更优的失真-真实感平衡。

通过优化率失真-真实感权衡,生成式图像压缩方法能产出比传统率失真优化模型更细腻、更真实的重建图像。本文提出一种注入扩散知识的新型深度学习生成式图像压缩方法,在实际场景中显著提升纹理还原能力。从三个维度探索该任务中的率失真-真实感权衡:首先,认识到图像纹理与频域特性的强关联,设计了分形频域感知带压缩(FFAB-IC)网络,有效捕获自然图像中的方向性频率成分;该网络在神经非线性映射中集成常用分形带特征操作,增强保留关键信息与过滤冗余细节的能力。其次,为在有限带宽下提升重建视觉质量,将扩散知识引入编码器,并在解码过程中实施扩散迭代,从而有效恢复丢失的纹理细节。最后,为充分利用空间与频域强度信息,引入频域与内容感知正则化项,指导生成式压缩网络训练。大量定量与定性实验表明,所提方法在可实现的失真-真实感组合边界上取得突破,即在高真实感下实现更优失真,或在低失真下实现更高真实感,优于以往所有方法。

原文摘要 · Abstract (English)

By optimizing the rate-distortion-realism trade-off, generative image compression approaches produce detailed, realistic images instead of the only sharp-looking reconstructions produced by rate-distortion-optimized models. In this paper, we propose a novel deep learning-based generative image compression method injected with diffusion knowledge, obtaining the capacity to recover more realistic textures in practical scenarios. Efforts are made from three perspectives to navigate the rate-distortion-realism trade-off in the generative image compression task. First, recognizing the strong connection between image texture and frequency-domain characteristics, we design a Fractal Frequency-Aware Band Image Compression (FFAB-IC) network to effectively capture the directional frequency components inherent in natural images. This network integrates commonly used fractal band feature operations within a neural non-linear mapping design, enhancing its ability to retain essential given information and filter out unnecessary details. Then, to improve the visual quality of image reconstruction under limited bandwidth, we integrate diffusion knowledge into the encoder and implement diffusion iterations into the decoder process, thus effectively recovering lost texture details. Finally, to fully leverage the spatial and frequency intensity information, we incorporate frequency- and content-aware regularization terms to regularize the training of the generative image compression network. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of the proposed method, advancing the boundaries of achievable distortion-realism pairs, i.e., our method achieves better distortions at high realism and better realism at low distortion than ever before.

图像压缩扩散模型纹理还原

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。