arXiv:2608.10952cs.CV2026-08中稿 · ICIP 2026被引 1

用多尺度隐变量提升图像压缩效率,显著优于现有方法。

Multiple Scale Latents for Learned Image Compression

  • 采用多尺度隐变量与独立熵模型,更好捕捉图像结构
  • 在Kodak数据集上比VVC降低17.9%的码率(BD-rate)
  • 可兼容现有压缩技术,适合作为通用优化模块

大多数学习型图像压缩系统依赖单一隐变量与超先验,难以高效捕捉跨尺度的图像结构。本文提出分层隐变量表示,通过在不同尺度使用多个隐变量及各自独立的熵模型,更准确地建模隐变量的空间结构。实验表明,该方法在Kodak数据集上相比VVC实现17.9%的BD-rate降低,验证了多尺度隐变量表示的有效性。此外,该方法与其他学习压缩进展具有正交性,可作为通用组件无缝集成到现有框架中。

原文摘要 · Abstract (English)

Most learned image compression systems rely on a single latent representation combined with a hyperprior, which limits their ability to efficiently capture image structure across spatial scales. In this work, we propose a hierarchical latent representation to improve the efficiency of the entropy model. By using multiple latents at different scales, each with its own entropy model, we better capture the spatial structure of the latent representation. Our experiments show that this approach achieves a 17.9% BD-rate reduction over VVC on Kodak, demonstrating the effectiveness of multi-scale latent representations. Furthermore, the approach is orthogonal to other advances in learned image compression, making it a versatile addition to existing methods.

图像压缩隐变量多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。