用两种不同隐空间协同解码,实现图像压缩保真与感知真实性的平衡。
Dual-Latent Collaborative Decoding for Fidelity-Perception Balanced Image Compression

- 将连续隐空间和离散隐空间分别用于保真度和感知质量,分工协作。
- 在低比特率下相比基线提升感知质量1.2分(在LPIPS上降低0.13)。
- 适合需要兼顾视觉质量和压缩效率的应用,如视频流、移动端图像传输。
学习型图像压缩(LIC)需在多种比特率下平衡失真保真度与感知真实性。现有方法通常依赖单一压缩隐表示同时承载结构细节、语义线索与感知先验,导致角色冲突。连续的标量量化(SQ)隐空间支持率可调保真度,但低比特率下易丢失感知细节;离散的向量量化(VQ)隐空间保留紧凑语义信息,却受限于结构保真度与比特率扩展性。为此,我们提出混合解码专家(MoDE),一种双隐空间协同解码框架,将重建职责分配给互补的隐空间范式:将SQ分支视为保真导向专家,VQ分支视为感知导向专家,并通过两个解码端模块协调:专家特定增强(ESE)保留各分支专属参考,交叉专家调制(CEM)实现重建中选择性互补传递。该框架支持共享双流比特流下的选择性跨隐空间协作,实现保真锚定与感知锚定解码。大量实验表明,MoDE在广泛比特率范围内优于代表性失真导向、感知导向、生成式及双隐空间基线,验证了解码端专家协作作为广域保真-感知平衡压缩的有效设计。
原文摘要 · Abstract (English)
Learned image compression (LIC) increasingly requires reconstructions that balance distortion fidelity and perceptual realism across a wide range of bitrates. However, most existing methods still rely on a single compressed latent representation to simultaneously carry structural details, semantic cues, and perceptual priors, requiring the same latent representation to serve multiple, potentially conflicting roles. This tension becomes evident across different latent paradigms: scalar-quantized (SQ) continuous latents provide rate-scalable fidelity but tend to lose perceptual details at low rates, while vector-quantized (VQ) discrete tokens preserve compact semantic cues but suffer from limited structural fidelity and bitrate scalability. To address this issue, we propose Mixture of Decoder Experts (MoDE), a dual-latent collaborative decoding framework that decomposes reconstruction responsibilities across complementary latent paradigms. Specifically, MoDE treats the SQ branch as a fidelity-oriented expert and the VQ branch as a perception-oriented expert, and coordinates them through two decoder-side modules: Expert-Specific Enhancement (ESE), which preserves branch-specific expert references, and Cross-Expert Modulation (CEM), which enables selective complementary transfer during reconstruction. The resulting framework supports selective cross-latent collaboration under a shared dual-stream bitstream and enables both fidelity-anchored and perception-anchored decoding. Extensive experiments demonstrate that MoDE achieves a more favorable fidelity-perception balance than representative distortion-oriented, perception-oriented, generative, and dual-latent baselines across a wide bitrate range, highlighting decoder-side expert collaboration as an effective design for wide-range fidelity-perception balanced LIC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。