arXiv:2503.21284cs.CVcs.AI2025-03中稿 · publication in IEE…被引 7

用可逆网络实现高保真图像压缩,单模型覆盖全码率范围

Multi-Scale Invertible Neural Network for Wide-Range Variable-Rate Learned Image Compression

  • 基于多尺度可逆神经网络,实现输入图像到多级潜在表示的双向映射
  • 在全码率范围内超越VVC标准,尤其在高码率下显著提升压缩质量
  • 适合需要统一模型支持多种码率的实用图像压缩场景

基于自编码器的结构主导了近期学习型图像压缩方法,但其固有的信息损失限制了高码率下的率失真性能,并制约了码率灵活调整能力。本文提出一种基于可逆变换的变码率图像压缩模型以克服上述局限。具体地,设计了一种轻量级多尺度可逆神经网络,将输入图像双射映射为多尺度潜在表示。为提升压缩效率,引入带有扩展增益单元的多尺度空间-通道上下文模型,从高到低层级估计潜在表示的熵。实验表明,所提方法在现有变码率方法中达到领先性能,且与最新多模型方法相比仍具竞争力。尤为关键的是,该方法是首个在极宽码率范围内仅用单一模型即超越VVC标准的学习除了图像压缩方案,尤其在高码率下表现突出。源代码已开源:https://github.com/hytu99/MSINN-VRLIC。

原文摘要 · Abstract (English)

Autoencoder-based structures have dominated recent learned image compression methods. However, the inherent information loss associated with autoencoders limits their rate-distortion performance at high bit rates and restricts their flexibility of rate adaptation. In this paper, we present a variable-rate image compression model based on invertible transform to overcome these limitations. Specifically, we design a lightweight multi-scale invertible neural network, which bijectively maps the input image into multi-scale latent representations. To improve the compression efficiency, a multi-scale spatial-channel context model with extended gain units is devised to estimate the entropy of the latent representation from high to low levels. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared to existing variable-rate methods, and remains competitive with recent multi-model approaches. Notably, our method is the first learned image compression solution that outperforms VVC across a very wide range of bit rates using a single model, especially at high bit rates. The source code is available at https://github.com/hytu99/MSINN-VRLIC.

图像压缩可逆网络变码率深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。