提出自适应结构的神经视频压缩方法,提升动态内容捕捉能力。
CANeRV: Content Adaptive Neural Representation for Video Compression
- 根据视频内容动态调整网络结构,实现帧间与序列级自适应优化。
- 在多个数据集上超越H.266/VVC和现有神经压缩方法,压缩性能更优。
- 适合需要高画质低码率视频编码的应用场景,如流媒体传输。
近年来,基于隐式神经表示(INR)的视频压缩方法通过全局优化网络参数,有效捕捉整个视频序列的依赖关系与特征,展现出优越的压缩潜力。然而,多数现有INR方法采用固定统一的网络结构,难以适应视频内容在时序上的动态变化,导致压缩效果受限。为此,本文提出内容自适应神经表示视频压缩方法(CANeRV),通过动态序列级调整(DSA)增强跨序列动态信息捕捉,利用动态帧级调整(DFA)提升帧间动态建模能力,并设计层次化结构自适应机制(HSA)以更好保留视频帧内空间结构细节,显著提升细节重建能力。实验表明,CANeRV在多个视频数据集上均优于H.266/VVC及当前最先进的INR基视频压缩方法。
原文摘要 · Abstract (English)
Recent advances in video compression introduce implicit neural representation (INR) based methods, which effectively capture global dependencies and characteristics of entire video sequences. Unlike traditional and deep learning based approaches, INR-based methods optimize network parameters from a global perspective, resulting in superior compression potential. However, most current INR methods utilize a fixed and uniform network architecture across all frames, limiting their adaptability to dynamic variations within and between video sequences. This often leads to suboptimal compression outcomes as these methods struggle to capture the distinct nuances and transitions in video content. To overcome these challenges, we propose Content Adaptive Neural Representation for Video Compression (CANeRV), an innovative INR-based video compression network that adaptively conducts structure optimisation based on the specific content of each video sequence. To better capture dynamic information across video sequences, we propose a dynamic sequence-level adjustment (DSA). Furthermore, to enhance the capture of dynamics between frames within a sequence, we implement a dynamic frame-level adjustment (DFA). {Finally, to effectively capture spatial structural information within video frames, thereby enhancing the detail restoration capabilities of CANeRV, we devise a structure level hierarchical structural adaptation (HSA).} Experimental results demonstrate that CANeRV can outperform both H.266/VVC and state-of-the-art INR-based video compression techniques across diverse video datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。