用可学习字典提升图像压缩熵模型性能
Learned Image Compression with Dictionary-based Entropy Model
- 引入可学习字典捕捉训练数据中的典型结构
- 在多个基准数据集上达到顶尖压缩效果
- 兼顾压缩性能与计算延迟,适合实际应用
学习型图像压缩方法近年来受到广泛关注,其率失真性能已超越现有最优传统压缩标准。熵模型在其中起关键作用,用于估计潜在表示的概率分布以进行熵编码。现有方法多采用超先验和自回归架构,仅关注潜在表示内部依赖关系,忽视从训练数据中提取先验信息的重要性。本文提出一种新型熵模型——基于字典的交叉注意力熵模型,通过引入可学习字典来总结训练数据中的典型结构,从而增强熵模型建模能力。大量实验表明,该模型在性能与延迟之间取得更优平衡,在多个基准数据集上实现当前最佳结果。
原文摘要 · Abstract (English)
Learned image compression methods have attracted great research interest and exhibited superior rate-distortion performance to the best classical image compression standards of the present. The entropy model plays a key role in learned image compression, which estimates the probability distribution of the latent representation for further entropy coding. Most existing methods employed hyper-prior and auto-regressive architectures to form their entropy models. However, they only aimed to explore the internal dependencies of latent representation while neglecting the importance of extracting prior from training data. In this work, we propose a novel entropy model named Dictionary-based Cross Attention Entropy model, which introduces a learnable dictionary to summarize the typical structures occurring in the training dataset to enhance the entropy model. Extensive experimental results have demonstrated that the proposed model strikes a better balance between performance and latency, achieving state-of-the-art results on various benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。