arXiv:2504.19119eess.IV2025-04中稿 · ACM TOMM被引 8

MLICv2提升图像压缩性能,解决高码率退化与全局相关性建模不足问题。

MLICv2: Enhanced Multi-Reference Entropy Modeling for Learned Image Compression

  • 引入轻量级令牌混合模块增强变换能力,缓解高码率性能下降
  • 通过超先验引导的全局上下文预测与通道重加权,提升熵建模精度
  • 支持实例特异性优化,适合追求极致压缩效率的研究者

近年来,学习型图像压缩(LIC)在性能上显著超越传统编码器。其中,基于多参考熵模型的MLIC系列已大幅领先于如内嵌式通用视频编码(VVC Intra)等传统方案。然而现有MLIC变体仍存在若干局限:高码率下因变换能力不足导致性能下降,初始片层熵建模无法捕捉全局相关性,且缺乏自适应通道重要性建模。本文提出MLICv2与MLICv2+,通过改进变换设计、先进熵建模及探索实例特异性优化潜力,系统性解决上述问题。变换方面,引入受MetaFormer启发的轻量级令牌混合块,有效缓解高码率性能退化同时保持计算效率;熵建模方面,提出超先验引导的全局相关性预测,实现初始片层中全局上下文提取,并引入通道重加权模块动态强化信息通道;进一步探索增强位置嵌入与引导选择性压缩策略以优化上下文建模。此外,采用随机Gumbel退火(SGA)方法验证输入特异性优化带来的性能提升潜力。大量实验表明,与VTM-17.0 Intra相比,MLICv2和MLICv2+在Kodak、Tecnick和CLIC Pro Val数据集上分别实现了Bjøntegaard-Delta Rate降低16.54%、21.61%、16.05%,以及20.46%、24.35%、19.14%。

原文摘要 · Abstract (English)

Recent advances in learned image compression (LIC) have achieved remarkable performance improvements over traditional codecs. Notably, the MLIC series-LICs equipped with multi-reference entropy models-have substantially surpassed conventional image codecs such as Versatile Video Coding (VVC) Intra. However, existing MLIC variants suffer from several limitations: performance degradation at high bitrates due to insufficient transform capacity, suboptimal entropy modeling that fails to capture global correlations in initial slices, and lack of adaptive channel importance modeling. In this paper, we propose MLICv2 and MLICv2+, enhanced successors that systematically address these limitations through improved transform design, dvanced entropy modeling, and exploration of the potential of instance-specific optimization. For transform enhancement, we introduce a lightweight token mixing block inspired by the MetaFormer architecture, which effectively mitigates high-bitrate performance degradation while maintaining computational efficiency. For entropy modeling improvements, we propose hyperprior-guided global correlation prediction to extract global context even in the initial slice of latent representation, complemented by a channel reweighting module that dynamically emphasizes informative channels. We further explore enhanced positional embedding and guided selective compression strategies for superior context modeling. Additionally, we apply the Stochastic Gumbel Annealing (SGA) to demonstrate the potential for further performance improvements through input-specific optimization. Extensive experiments demonstrate that MLICv2 and MLICv2+ achieve state-of-the-art results, reducing Bjøntegaard-Delta Rate by 16.54%, 21.61%, 16.05% and 20.46%, 24.35%, 19.14% on Kodak, Tecnick, and CLIC Pro Val datasets, respectively, compared to VTM-17.0 Intra.

图像压缩熵建模深度学习自适应优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。