arXiv:2505.24015eess.IV2025-05被引 2

用语义引导生成图像压缩,提升画质并加快速度。

Semantics-Guided Generative Image Compression

  • 引入语义分割指导生成解码,提升图像细节还原。
  • 内容自适应扩散减少推理步骤,编码解码提速超36%。
  • 在感知质量上超越主流压缩算法,适合高保真图像传输。

大型多模态模型在文本到图像生成方面的进展正拓展至图像压缩领域,实现了极低比特率下的高质量图像表示。本文针对现有多模态图像语义压缩(MISC)方法,提出新组件:利用语义分割指导生成解码器,并引入内容自适应扩散机制,根据图像特征动态调整扩散步数。实验结果表明,所提方法显著提升基线MISC模型的性能,同时降低计算复杂度,编码与解码时间均减少超过36%。此外,该压缩框架在感知相似性和图像质量方面优于主流编解码器。代码与可视化示例已公开。

原文摘要 · Abstract (English)

Advancements in text-to-image generative AI with large multimodal models are spreading into the field of image compression, creating high-quality representation of images at extremely low bit rates. This work introduces novel components to the existing multimodal image semantic compression (MISC) approach, enhancing the quality of the generated images in terms of PSNR and perceptual metrics. The new components include semantic segmentation guidance for the generative decoder, as well as content-adaptive diffusion, which controls the number of diffusion steps based on image characteristics. The results show that our newly introduced methods significantly improve the baseline MISC model while also decreasing the complexity. As a result, both the encoding and decoding time are reduced by more than 36%. Moreover, the proposed compression framework outperforms mainstream codecs in terms of perceptual similarity and quality. The code and visual examples are available.

图像压缩生成模型语义引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。