通过多样化上下文提升神经图像编码熵模型性能,速度更快、压缩更好。
Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec
- 引入多种超隐变量表示,增强前向建模的上下文信息。
- 在各码率下均实现性能提升,Kodak数据集上比最优基线低3.73% BD-rate。
- 适合追求高效高保真图像压缩的研究与工程人员。
设计快速高效的熵模型对神经编码器的实际应用至关重要。除空间自回归模型外,基于反向自适应的更高效模型近年来被提出,它们通过减少建模步骤降低解码时间,同时借助更丰富的上下文提升或维持率失真性能。然而,现有方法受限于沿用前向自适应的设计惯例:仅使用单一类型的超隐变量表示,尤其在首步建模时上下文信息不足。本文提出一种简单而有效的熵建模框架,在不增加比特率的前提下,为前向自适应提供充分上下文。具体而言,引入多样化超隐变量表示,即在原有基础上增加两种新类型上下文;并提出有效利用这些多样上下文进行元素上下文化的策略。实验结果表明,该框架在多个主流数据集上一致提升率失真性能,例如在Kodak数据集上相比当前最优基线实现3.73%的BD-rate增益。
原文摘要 · Abstract (English)
Designing a fast and effective entropy model is challenging but essential for practical application of neural codecs. Beyond spatial autoregressive entropy models, more efficient backward adaptation-based entropy models have been recently developed. They not only reduce decoding time by using smaller number of modeling steps but also maintain or even improve rate--distortion performance by leveraging more diverse contexts for backward adaptation. Despite their significant progress, we argue that their performance has been limited by the simple adoption of the design convention for forward adaptation: using only a single type of hyper latent representation, which does not provide sufficient contextual information, especially in the first modeling step. In this paper, we propose a simple yet effective entropy modeling framework that leverages sufficient contexts for forward adaptation without compromising on bit-rate. Specifically, we introduce a strategy of diversifying hyper latent representations for forward adaptation, i.e., using two additional types of contexts along with the existing single type of context. In addition, we present a method to effectively use the diverse contexts for contextualizing the current elements to be encoded/decoded. By addressing the limitation of the previous approach, our proposed framework leads to significant performance improvements. Experimental results on popular datasets show that our proposed framework consistently improves rate--distortion performance across various bit-rate regions, e.g., 3.73% BD-rate gain over the state-of-the-art baseline on the Kodak dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。