arXiv:2604.06564eess.IVcs.CV2026-04中稿 · IEEE Transactions …

用混合神经网络与残差网格建模视频,提升细节表现力和重建质量。

CWRNN-INVR: A Coupled WarpRNN based Implicit Neural Video Representation

  • 分治式设计:神经网络处理规律结构,残差网格捕捉不规则细节
  • 耦合WarpRNN实现多尺度运动建模,提升动态信息表达能力
  • 适用于高质量视频重建与下游任务,代码开源可复现

隐式神经视频表征(INVR)作为一种新兴的视频表示与压缩方法,利用可学习的网格与神经网络进行建模。现有方法侧重于改进网格结构与具备强表征能力的神经网络架构,但未深入探讨二者在视频表征中的作用差异。本文从视频信息构成角度首次分析基于神经网络与基于网格的INVR的差异,揭示其各自优势:神经网络适合表征通用结构,网格更擅长捕捉具体细节。为此,提出一种融合神经网络与残差网格的框架,其中神经网络用于表征视频中的规律性结构信息,残差网格则用于表示剩余的非规律性信息。设计了基于耦合WarpRNN的多尺度运动表示与补偿模块,显式建模规律性运动特征,因此命名为CWRNN-INVR。对于不规则信息,通过联合学习不规则外观与运动的混合残差网格实现表征。该混合残差网格可与耦合WarpRNN协同工作,支持网络复用。实验表明,本方法在UVG数据集上以3M模型达到平均33.73 dB的PSNR,优于现有INVR方法,并在其他下游任务中表现更优。代码已开源。

原文摘要 · Abstract (English)

Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent representation and neural network architectures with large representation capability, lacking the study on their roles in video representation. In this paper, the difference between INVR based on neural network and INVR based on grid is first investigated from the perspective of video information composition to specify their own advantages, i.e., neural network for general structure while grid for specific detail. Accordingly, an INVR based on mixed neural network and residual grid framework is proposed, where the neural network is used to represent the regular and structured information and the residual grid is used to represent the remaining irregular information in a video. A Coupled WarpRNN-based multi-scale motion representation and compensation module is specifically designed to explicitly represent the regular and structured information, thus terming our method as CWRNN-INVR. For the irregular information, a mixed residual grid is learned where the irregular appearance and motion information are represented together. The mixed residual grid can be combined with the coupled WarpRNN in a way that allows for network reuse. Experiments show that our method achieves the best reconstruction results compared with the existing methods, with an average PSNR of 33.73 dB on the UVG dataset under the 3M model and outperforms existing INVR methods in other downstream tasks. The code can be found at https://github.com/yiyang-sdu/CWRNN-INVR.git}{https://github.com/yiyang-sdu/CWRNN-INVR.git.

视频表征隐式表示运动建模神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。