arXiv:2502.07160cs.CVcs.MM2025-02中稿 · PRICAI 2025被引 7

融合扩散模型与传统压缩,实现超低码率下高保真图像还原。

HDCompression: Hybrid-Diffusion Image Compression for Ultra-Low Bitrates

  • 双流架构结合扩散模型与传统压缩,互补提升质量。
  • 在0.01比特/像素下仍保持高视觉质量,远超现有方法。
  • 适合对压缩后图像质量要求极高的场景,如医疗或卫星图像。

在超低码率下,传统学习图像压缩(LIC)因严重量化导致明显伪影,而生成式向量量化(VQ)建模则因生成先验与输入不匹配造成保真度下降。本文提出混合扩散图像压缩(HDCompression),采用双流框架融合生成式VQ建模、扩散模型与传统LIC,兼顾高保真与高感知质量。不同于以往直接用预训练LIC生成低质量信息的方法,本工作利用扩散模型从真实图像中提取高质量互补保真信息,有效提升索引图预测、增强LIC流的保真输出,并通过VQ隐变量修正优化条件重构。扩散模型基于轻量级密集代表向量(DRV),采样调度简单。大量实验表明,HDCompression在定量指标和定性可视化上均优于现有传统LIC、生成式VQ及混合框架,在超低码率下表现出均衡且鲁棒的压缩性能。

原文摘要 · Abstract (English)

Image compression under ultra-low bitrates remains challenging for both conventional learned image compression (LIC) and generative vector-quantized (VQ) modeling. Conventional LIC suffers from severe artifacts due to heavy quantization, while generative VQ modeling gives poor fidelity due to the mismatch between learned generative priors and specific inputs. In this work, we propose Hybrid-Diffusion Image Compression (HDCompression), a dual-stream framework that utilizes both generative VQ-modeling and diffusion models, as well as conventional LIC, to achieve both high fidelity and high perceptual quality. Different from previous hybrid methods that directly use pre-trained LIC models to generate low-quality fidelity-preserving information from heavily quantized latent, we use diffusion models to extract high-quality complementary fidelity information from the ground-truth input, which can enhance the system performance in several aspects: improving index map prediction, enhancing the fidelity-preserving output of the LIC stream, and refining conditioned image reconstruction with VQ-latent correction. In addition, our diffusion model is based on a dense representative vector (DRV), which is lightweight with very simple sampling schedulers. Extensive experiments demonstrate that our HDCompression outperforms the previous conventional LIC, generative VQ-modeling, and hybrid frameworks in both quantitative metrics and qualitative visualization, providing balanced robust compression performance at ultra-low bitrates.

图像压缩扩散模型超低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。