arXiv:2409.07414cs.CVeess.IV2024-09NeurIPS被引 37

用神经网络压缩视频表示,实现端到端优化,性能超VVC 24%。

NVRC: Neural Video Representation Compression

  • 设计端到端可训练的神经视频表示压缩框架
  • 在UVG数据集上比VVC VTM提升24% PSNR
  • 适合关注神经编码与视频压缩前沿的研究者

基于隐式神经表示(INR)的视频编码近年来展现出与传统及学习型方法竞争的潜力。现有方法通过训练神经网络过拟合视频序列,并压缩其参数以获得紧凑表示。然而,尽管取得进展,最佳的INR方法仍落后于最新标准编码器(如VVC VTM),部分原因在于采用的模型压缩技术较为简单。本文提出新型INR视频压缩框架NVRC,聚焦于表示本身的压缩。基于新提出的熵编码与量化模型,NVRC首次实现端到端可优化的INR视频编码。为降低熵模型引入的额外码率开销,我们还提出一种分层压缩框架,用于编码网络、量化及熵模型的所有参数。实验表明,NVRC在UVG数据集上相比VVC VTM(随机访问模式)平均提升24% PSNR,据我们所知,这是首个达到此性能的INR视频编码方案。NVRC的实现将公开发布。

原文摘要 · Abstract (English)

Recent advances in implicit neural representation (INR)-based video coding have demonstrated its potential to compete with both conventional and other learning-based approaches. With INR methods, a neural network is trained to overfit a video sequence, with its parameters compressed to obtain a compact representation of the video content. However, although promising results have been achieved, the best INR-based methods are still out-performed by the latest standard codecs, such as VVC VTM, partially due to the simple model compression techniques employed. In this paper, rather than focusing on representation architectures as in many existing works, we propose a novel INR-based video compression framework, Neural Video Representation Compression (NVRC), targeting compression of the representation. Based on the novel entropy coding and quantization models proposed, NVRC, for the first time, is able to optimize an INR-based video codec in a fully end-to-end manner. To further minimize the additional bitrate overhead introduced by the entropy models, we have also proposed a new model compression framework for coding all the network, quantization and entropy model parameters hierarchically. Our experiments show that NVRC outperforms many conventional and learning-based benchmark codecs, with a 24% average coding gain over VVC VTM (Random Access) on the UVG dataset, measured in PSNR. As far as we are aware, this is the first time an INR-based video codec achieving such performance. The implementation of NVRC will be released.

视频压缩神经表示端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。