用频率感知机制加速扩散模型图像压缩,实现更高效高保真重建。
Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement
- 引入频率解耦注意力,动态优化扩散过程中的噪声预测
- 两步解码实现超10倍速度提升,比特率降低65.1%(PSNR)
- 无需重训练主模型,适合部署于资源受限的低码率场景
基于扩散模型的生成先验在极低码率下已实现视觉上令人信服的图像压缩。然而,现有方法因训练范式碎片化,存在采样缓慢和比特分配不佳的问题。本文提出一种名为DiffCR的新压缩框架,通过频率感知跳过估计(FaSE)模块,对预训练潜空间扩散模型的ε-预测先验进行精炼,并利用频率解耦注意力(FDA)在不同时间步对齐压缩潜变量。此外,轻量级一致性估计器支持快速两步解码,保持扩散采样的语义轨迹。无需更新主扩散模型,DiffCR相较当前最优扩散基压缩方法实现27.2%(LPIPS)和65.1%(PSNR)的BD-rate降低,且解码速度提升超过10倍。
原文摘要 · Abstract (English)
Recent advancements in diffusion-based generative priors have enabled visually plausible image compression at extremely low bit rates. However, existing approaches suffer from slow sampling processes and suboptimal bit allocation due to fragmented training paradigms. In this work, we propose Accelerate \textbf{Diff}usion-based Image Compression via \textbf{C}onsistency Prior \textbf{R}efinement (DiffCR), a novel compression framework for efficient and high-fidelity image reconstruction. At the heart of DiffCR is a Frequency-aware Skip Estimation (FaSE) module that refines the $ε$-prediction prior from a pre-trained latent diffusion model and aligns it with compressed latents at different timesteps via Frequency Decoupling Attention (FDA). Furthermore, a lightweight consistency estimator enables fast \textbf{two-step decoding} by preserving the semantic trajectory of diffusion sampling. Without updating the backbone diffusion model, DiffCR achieves substantial bitrate savings (27.2\% BD-rate (LPIPS) and 65.1\% BD-rate (PSNR)) and over $10\times$ speed-up compared to SOTA diffusion-based compression baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。