多尺度融合提升神经视频表示,动态内容压缩更优
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
- 用多尺度空间解码器和自适应损失融合多分辨率信息
- 在HEVC ClassB和UVG数据集上优于现有神经编码方法
- 适合需要高动态内容压缩的视频应用
隐式神经表示(INRs)在视频压缩领域展现出潜力,已达到与最新编码标准如H.266/VVC相当的性能。然而,现有基于INR的方法在细节丰富且快速变化的视频内容上表现不佳,主要源于对网络内部特征利用不足及缺乏视频特异性设计。为此,我们提出多尺度特征融合框架MSNeRV。编码阶段采用时间窗口划分视频为多个图像组(GoPs),GoP级网格用于背景建模,并设计多尺度空间解码器与尺度自适应损失函数,整合多分辨率和多频域信息。进一步引入多尺度特征块,充分挖掘隐藏特征。在HEVC ClassB和UVG数据集上的实验表明,该模型在INR方法中具备更强表示能力,在动态场景下压缩效率超过VTM-23.7(随机访问模式)。
原文摘要 · Abstract (English)
Implicit Neural representations (INRs) have emerged as a promising approach for video compression, and have achieved comparable performance to the state-of-the-art codecs such as H.266/VVC. However, existing INR-based methods struggle to effectively represent detail-intensive and fast-changing video content. This limitation mainly stems from the underutilization of internal network features and the absence of video-specific considerations in network design. To address these challenges, we propose a multi-scale feature fusion framework, MSNeRV, for neural video representation. In the encoding stage, we enhance temporal consistency by employing temporal windows, and divide the video into multiple Groups of Pictures (GoPs), where a GoP-level grid is used for background representation. Additionally, we design a multi-scale spatial decoder with a scale-adaptive loss function to integrate multi-resolution and multi-frequency information. To further improve feature extraction, we introduce a multi-scale feature block that fully leverages hidden features. We evaluate MSNeRV on HEVC ClassB and UVG datasets for video representation and compression. Experimental results demonstrate that our model exhibits superior representation capability among INR-based approaches and surpasses VTM-23.7 (Random Access) in dynamic scenarios in terms of compression efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。