arXiv:2511.08009eess.IVcs.CV2025-11

用高斯噪声生成图像专属潜在表示,无需传输编码,压缩效果佳。

From Noise to Latent: Generating Gaussian Latents for INR-Based Image Compression

  • 从高斯噪声直接生成多尺度潜在表示,用共享随机种子确定性生成。
  • 在Kodak和CLIC数据集上达到与现有方法相当的率失真性能。
  • 适合追求低传输开销且希望保留潜在表示优势的研究者。

基于隐式神经表示(INR)的图像压缩方法通过过拟合图像特定的潜在编码已展现出竞争力,但由于缺乏表达性强的潜在表示,仍不及端到端(E2E)压缩方法。而E2E方法依赖传输潜在编码并使用复杂的熵模型,导致解码复杂度增加。受E2E编码器中将潜在表示转化为高斯噪声以消除空间冗余的启发,本文探索逆向路径:直接从高斯噪声生成潜在表示。提出一种新压缩范式,通过共享随机种子确定性地生成多尺度高斯噪声张量,并利用高斯参数预测(GPP)模块估计分布参数,结合重参数化技巧实现一次性潜在生成。生成的潜在表示经合成网络重建图像。该方法无需传输潜在编码,同时保留了潜在表示的优势,在Kodak和CLIC数据集上实现了具有竞争力的率失真表现。据我们所知,这是首个探索高斯潜在生成用于学习型图像压缩的工作。

原文摘要 · Abstract (English)

Recent implicit neural representation (INR)-based image compression methods have shown competitive performance by overfitting image-specific latent codes. However, they remain inferior to end-to-end (E2E) compression approaches due to the absence of expressive latent representations. On the other hand, E2E methods rely on transmitting latent codes and requiring complex entropy models, leading to increased decoding complexity. Inspired by the normalization strategy in E2E codecs where latents are transformed into Gaussian noise to demonstrate the removal of spatial redundancy, we explore the inverse direction: generating latents directly from Gaussian noise. In this paper, we propose a novel image compression paradigm that reconstructs image-specific latents from a multi-scale Gaussian noise tensor, deterministically generated using a shared random seed. A Gaussian Parameter Prediction (GPP) module estimates the distribution parameters, enabling one-shot latent generation via reparameterization trick. The predicted latent is then passed through a synthesis network to reconstruct the image. Our method eliminates the need to transmit latent codes while preserving latent-based benefits, achieving competitive rate-distortion performance on Kodak and CLIC dataset. To the best of our knowledge, this is the first work to explore Gaussian latent generation for learned image compression.

图像压缩潜在表示高斯噪声INR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。