arXiv:2411.17160eess.IVcs.CV2024-11被引 1

无需运动估计的神经视频压缩,效率更高且更省资源。

Motion Free B-frame Coding for Neural Video Compression

  • 用基于核的编码器替代传统运动与残差双模块结构。
  • 在HEVC-B数据集上超越现有最优方法,模型体积小3-4倍。
  • 消除运动估计计算,提升画质并降低模糊伪影。

主流深度神经视频压缩网络通常采用经典视频编码的混合架构,包含独立的运动编码与残差编码模块,并普遍使用对称自编码器作为基础结构。本文提出一种新型无运动编码方案——基于核的无运动视频压缩。该方法通过消除运动估计算法、运动补偿及运动编码三大耗时模块,显著降低计算复杂度并提升编码效率。同时,基于核的自编码器缓解了传统对称自编码器常见的模糊伪影问题,有效改善重建帧视觉质量。实验表明,该框架在HEVC-class B数据集上优于当前最先进方法,在UVG和MCL-JCV数据集上表现相当,且模型尺寸仅为运动型网络的三至四分之一。

原文摘要 · Abstract (English)

Typical deep neural video compression networks usually follow the hybrid approach of classical video coding that contains two separate modules: motion coding and residual coding. In addition, a symmetric auto-encoder is often used as a normal architecture for both motion and residual coding. In this paper, we propose a novel approach that handles the drawbacks of the two typical above-mentioned architectures, we call it kernel-based motion-free video coding. The advantages of the motion-free approach are twofold: it improves the coding efficiency of the network and significantly reduces computational complexity thanks to eliminating motion estimation, motion compensation, and motion coding which are the most time-consuming engines. In addition, the kernel-based auto-encoder alleviates blur artifacts that usually occur with the conventional symmetric autoencoder. Consequently, it improves the visual quality of the reconstructed frames. Experimental results show the proposed framework outperforms the SOTA deep neural video compression networks on the HEVC-class B dataset and is competitive on the UVG and MCL-JCV datasets. In addition, it generates high-quality reconstructed frames in comparison with conventional motion coding-based symmetric auto-encoder meanwhile its model size is much smaller than that of the motion-based networks around three to four times.

视频压缩神经编码无运动估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。