arXiv:2503.01428cs.CVeess.IV2025-03ICCV被引 19

通过分离语义与细节,实现低于0.01 bpp的极致图像压缩

DLF: Extreme Image Compression with Dual-generative Latent Fusion

  • 将隐空间分解为语义和细节两分支,分别压缩
  • 在CLIC2020上比MS-ILLM节省27.93%比特,LPIPS降低53.55%
  • 适合追求极致压缩率且重视视觉保真的应用

近年来,极端图像压缩通过压缩生成式分词器的标记取得了显著进展。然而,这些方法常侧重于数据集内共性语义的聚类,忽视了单个物体的多样化细节,导致低比特率下重建保真度下降。为此,我们提出双生成隐空间融合(DLF)范式:将隐空间分解为语义与细节元素,通过两个独立分支分别压缩。语义分支将高层信息聚类为紧凑标记,细节分支编码感知关键细节以提升整体保真度。此外,设计跨分支交互机制减少冗余,降低总比特开销。实验表明,DLF在低于0.01比特每像素(bpp)时仍保持优异重建质量。在CLIC2020测试集上,相比MS-ILLM,LPIPS降低27.93%,DISTS降低53.55%。同时,其视觉保真度超越近期基于扩散模型的编码器,且生成真实感相当。

原文摘要 · Abstract (English)

Recent studies in extreme image compression have achieved remarkable performance by compressing the tokens from generative tokenizers. However, these methods often prioritize clustering common semantics within the dataset, while overlooking the diverse details of individual objects. Consequently, this results in suboptimal reconstruction fidelity, especially at low bitrates. To address this issue, we introduce a Dual-generative Latent Fusion (DLF) paradigm. DLF decomposes the latent into semantic and detail elements, compressing them through two distinct branches. The semantic branch clusters high-level information into compact tokens, while the detail branch encodes perceptually critical details to enhance the overall fidelity. Additionally, we propose a cross-branch interactive design to reduce redundancy between the two branches, thereby minimizing the overall bit cost. Experimental results demonstrate the impressive reconstruction quality of DLF even below 0.01 bits per pixel (bpp). On the CLIC2020 test set, our method achieves bitrate savings of up to 27.93% on LPIPS and 53.55% on DISTS compared to MS-ILLM. Furthermore, DLF surpasses recent diffusion-based codecs in visual fidelity while maintaining a comparable level of generative realism. Project: https://dlfcodec.github.io/

图像压缩生成模型极低码率隐空间融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。