arXiv:2409.06586eess.IV2024-09中稿 · EUSIPCO European c…被引 2

用可调输入尺度实现通用端到端图像压缩,速率失真更优。

Universal End-to-End Neural Network for Lossy Image Compression

  • 推理时仅调整输入尺度,无需改架构或损失函数。
  • 跨多种架构与训练方法表现稳定,适用性广。
  • 结构简单,降低计算与内存开销,适合实际部署。

本文提出一种基于变分自编码器(VAE)的可变码率有损图像压缩神经网络。通过在推理阶段仅调整输入尺度,实现高效率-失真权衡。在多种VAE压缩架构(如CNN、ViT)和训练策略(如MSE、SSIM)下均表现出卓越的通用性,归因于神经网络的内在泛化能力。相比需修改模型结构或损失函数的方法,本方案强调简洁性,显著降低计算复杂度与内存需求。实验不仅验证了方法的有效性,也展示了其推动可变码率神经网络图像压缩发展的潜力。

原文摘要 · Abstract (English)

This paper presents variable bitrate lossy image compression using a VAE-based neural network. An adaptable image quality adjustment strategy is proposed. The key innovation involves adeptly adjusting the input scale exclusively during the inference process, resulting in an exceptionally efficient rate-distortion mechanism. Through extensive experimentation, across diverse VAE-based compression architectures (CNN, ViT) and training methodologies (MSE, SSIM), our approach exhibits remarkable universality. This success is attributed to the inherent generalization capacity of neural networks. Unlike methods that adjust model architecture or loss functions, our approach emphasizes simplicity, reducing computational complexity and memory requirements. The experiments not only highlight the effectiveness of our approach but also indicate its potential to drive advancements in variable-rate neural network lossy image compression methodologies.

图像压缩变分自编码器端到端通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。