arXiv:2501.01231cs.CVcs.LG2025-01中稿 · IEEE TRANSACTIONS …

利用向量量化与熵梯度提升神经编解码器压缩效率

Exploiting Latent Properties to Optimize Neural Codecs

  • 采用预设最优均匀向量量化替代非均匀量化
  • 用解码端可获取的熵梯度代理重建误差梯度,提升性能
  • 适用于现有神经编解码器,显著降低码率1-3%

端到端图像和视频编解码器正日益超越传统依赖多年人工设计的压缩技术。这类可训练编解码器具备适应感知失真度量、在特定场景表现优异等优势。然而,当前最先进的神经编解码器尚未充分利用向量量化和解码设备中可用的熵梯度。本文提出利用这两项特性改进现成编解码器的性能:首先证明非均匀标量量化无法优于均匀量化,因此建议使用预设最优均匀向量量化;其次发现解码端的熵梯度与不可见的重建误差梯度相关,故将其作为代理以增强压缩性能。实验表明,该方法在多种预训练模型上可实现相同质量下码率降低1%至3%。此外,基于熵梯度的方案也显著提升了传统编解码器的表现。

原文摘要 · Abstract (English)

End-to-end image and video codecs are becoming increasingly competitive, compared to traditional compression techniques that have been developed through decades of manual engineering efforts. These trainable codecs have many advantages over traditional techniques, such as their straightforward adaptation to perceptual distortion metrics and high performance in specific fields thanks to their learning ability. However, current state-of-the-art neural codecs do not fully exploit the benefits of vector quantization and the existence of the entropy gradient in decoding devices. In this paper, we propose to leverage these two properties (vector quantization and entropy gradient) to improve the performance of off-the-shelf codecs. Firstly, we demonstrate that using non-uniform scalar quantization cannot improve performance over uniform quantization. We thus suggest using predefined optimal uniform vector quantization to improve performance. Secondly, we show that the entropy gradient, available at the decoder, is correlated with the reconstruction error gradient, which is not available at the decoder. We therefore use the former as a proxy to enhance compression performance. Our experimental results show that these approaches save between 1 to 3% of the rate for the same quality across various pretrained methods. In addition, the entropy gradient based solution improves traditional codec performance significantly as well.

神经编解码向量量化熵梯度压缩优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。