用显式语义+隐式纹理,实现超低码率图像压缩新突破
Dual-Representation Image Compression at Ultra-Low Bitrates via Explicit Semantics and Implicit Textures
- 显式语义与隐式纹理协同编码,无需训练
- 在超低码率下显著提升视觉质量,最高提升30%
- 可灵活调节失真与感知质量平衡,适合实际应用
尽管近期神经编解码器在低码率下优化感知质量表现良好,但在超低码率条件下性能急剧下降。为此,利用预训练模型语义先验的生成式压缩方法成为新方向。然而,现有方法受限于语义保真度与感知真实感之间的权衡:显式表示保留结构但缺乏细节,隐式方法虽能合成逼真纹理却导致语义偏移。本文提出一种统一框架,通过无训练方式协同整合显式与隐式表示。具体而言,以显式高层语义条件化扩散模型,并采用反向通道编码隐式传递细粒度信息。此外,引入即插即用编码器,通过调节隐式信息灵活控制失真-感知权衡。大量实验表明,该框架达到当前最优率-感知性能,在Kodak、DIV2K和CLIC2020数据集上,相比DiffC在DISTS BD-Rate上分别提升29.92%、19.33%和20.89%。
原文摘要 · Abstract (English)
While recent neural codecs achieve strong performance at low bitrates when optimized for perceptual quality, their effectiveness deteriorates significantly under ultra-low bitrate conditions. To mitigate this, generative compression methods leveraging semantic priors from pretrained models have emerged as a promising paradigm. However, existing approaches are fundamentally constrained by a tradeoff between semantic faithfulness and perceptual realism. Methods based on explicit representations preserve content structure but often lack fine-grained textures, whereas implicit methods can synthesize visually plausible details at the cost of semantic drift. In this work, we propose a unified framework that bridges this gap by coherently integrating explicit and implicit representations in a training-free manner. Specifically, We condition a diffusion model on explicit high-level semantics while employing reverse-channel coding to implicitly convey fine-grained details. Moreover, we introduce a plug-in encoder that enables flexible control of the distortion-perception tradeoff by modulating the implicit information. Extensive experiments demonstrate that the proposed framework achieves state-of-the-art rate-perception performance, outperforming existing methods and surpassing DiffC by 29.92%, 19.33%, and 20.89% in DISTS BD-Rate on the Kodak, DIV2K, and CLIC2020 datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。