用可调偏置建模视频运动,实现高效慢放与超分。
Bias for Action: Video Implicit Neural Representations with Bias Modulation
- 共享网络权重,每帧用独立偏置捕捉运动变化
- 支持10倍慢放、4倍空间超分辨率及去噪修复
- 适合需要连续视频建模的生成与编辑任务
我们提出一种基于隐式神经表示(INRs)的新连续视频建模框架ActINR。核心思想是将INRs视为可学习字典,其基函数形状由权重控制,位置由偏置决定。利用紧凑的非线性激活函数,我们假设偏置能有效捕获图像间的运动信息,从而实现视频序列的紧凑表示。ActINR在视频帧间共享INR权重,但为每帧设置唯一偏置,并将偏置建模为以时间索引为条件的独立INR输出,以保证平滑性。通过联合训练视频INR与偏置INR,我们实现了10倍视频慢放、4倍空间超分辨率结合2倍慢放、去噪与视频修复等能力。在多个视频处理任务中表现优异,性能提升常超过6dB,树立了连续视频建模新标准。
原文摘要 · Abstract (English)
We propose a new continuous video modeling framework based on implicit neural representations (INRs) called ActINR. At the core of our approach is the observation that INRs can be considered as a learnable dictionary, with the shapes of the basis functions governed by the weights of the INR, and their locations governed by the biases. Given compact non-linear activation functions, we hypothesize that an INR's biases are suitable to capture motion across images, and facilitate compact representations for video sequences. Using these observations, we design ActINR to share INR weights across frames of a video sequence, while using unique biases for each frame. We further model the biases as the output of a separate INR conditioned on time index to promote smoothness. By training the video INR and this bias INR together, we demonstrate unique capabilities, including $10\times$ video slow motion, 4x spatial super resolution along with 2x slow motion, denoising, and video inpainting. ActINR performs remarkably well across numerous video processing tasks (often achieving more than 6dB improvement), setting a new standard for continuous modeling of videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。