用分层时序神经表示提升视频压缩,更好捕捉运动细节与长程依赖。
Video Compression with Hierarchical Temporal Neural Representation
- 分层设计:融合相邻帧特征并按图片组自适应调参。
- 在相同码率下,重建质量优于现有基于INR的压缩方法。
- 适合需要高画质且低延迟的视频压缩场景。
视频压缩近年来受益于隐式神经表示(INRs),其将视频建模为连续函数,具备存储紧凑和重建灵活的优点,是传统编码器的有前景替代方案。然而,现有大多数基于INR的方法将时间维度作为独立输入,难以捕捉复杂的时间依赖性。为此,我们提出面向视频的分层时序神经表示TeNeRV。TeNeRV通过两个关键组件整合短时与长时依赖:首先,帧间特征融合(IFF)模块聚合邻近帧特征,强化局部时序一致性并捕捉精细运动;其次,图片组自适应调制(GAM)机制将视频划分为图片组(GoPs),学习每组特定先验,并调制网络参数,实现跨不同GoP的自适应表示。大量实验表明,TeNeRV在码率-失真性能上持续优于现有基于INR的方法,验证了所提方法的有效性。
原文摘要 · Abstract (English)
Video compression has recently benefited from implicit neural representations (INRs), which model videos as continuous functions. INRs offer compact storage and flexible reconstruction, providing a promising alternative to traditional codecs. However, most existing INR-based methods treat the temporal dimension as an independent input, limiting their ability to capture complex temporal dependencies. To address this, we propose a Hierarchical Temporal Neural Representation for Videos, TeNeRV. TeNeRV integrates short- and long-term dependencies through two key components. First, an Inter-Frame Feature Fusion (IFF) module aggregates features from adjacent frames, enforcing local temporal coherence and capturing fine-grained motion. Second, a GoP-Adaptive Modulation (GAM) mechanism partitions videos into Groups-of-Pictures and learns group-specific priors. The mechanism modulates network parameters, enabling adaptive representations across different GoPs. Extensive experiments demonstrate that TeNeRV consistently outperforms existing INR-based methods in rate-distortion performance, validating the effectiveness of our proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。