arXiv:2505.16687cs.CVeess.IV2025-05NeurIPS被引 23

一拍即合:用单步扩散实现超快高清图像压缩

One-Step Diffusion-Based Image Compression with Semantic Distillation

  • 单步扩散生成+潜空间压缩,跳过多步采样延迟
  • 超前1.4倍感知质量,码率降低39%,解码快20倍
  • 用隐式语义蒸馏替代文本提示,适合高保真压缩场景

尽管基于扩散的生成式图像编码器表现优异,但其迭代采样过程导致显著延迟。本文重新审视扩散编码器设计,认为生成压缩无需多步采样。提出OneDC——一种结合潜空间压缩模块与单步扩散生成器的新型编码器。针对单步扩散中语义引导的关键作用,提出以超先验作为语义信号,克服文本提示对复杂视觉内容表达能力的局限。为进一步增强超先验的语义能力,引入从预训练生成分词器迁移知识的语义蒸馏机制。同时采用像素域与潜空间联合优化策略,兼顾重建保真度与感知真实感。大量实验表明,OneDC在单步生成下达到顶尖感知质量,相比以往多步扩散编码器,码率降低39%以上,解码速度提升20倍。

原文摘要 · Abstract (English)

While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasing latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec -- that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20x faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/

图像压缩扩散模型单步生成语义蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。