用运动引导生成中间帧,无需训练就能提升视频插帧质量。
SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation

- 用对称非线性运动模型生成带对应关系的初始帧作为先验。
- 在预训练视频扩散模型中迭代引导,保持运动一致性与真实感。
- 适合需要高质量、高时序连贯性的视频处理应用。
我们提出对称非线性运动引导生成式视频帧插值(SNM-VFI),一种无需训练的运动可控生成式视频插帧框架,结合预训练光流模型与视频扩散模型。不同于传统基于扩散的方法从随机噪声生成中间帧,SNM-VFI利用对称非线性运动模型生成带对应关系的中间帧和置信图。这些光流引导帧被编码为潜在先验,用于初始化并迭代引导预训练视频扩散模型,使扩散过程在保持密集运动对应关系的同时提升感知真实感。为进一步提升输出质量,采用置信图将结构可靠的光流预测与扩散生成的细节在遮挡和物体边界等不确定区域进行融合。在DAVIS、Sintel和KITTI等挑战性基准上的大量实验表明,SNM-VFI在多种运动场景下实现了优异的感知质量、具有竞争力的重建精度和鲁棒的时序连贯性。
原文摘要 · Abstract (English)
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model. Specifically, we first utilize a pre-trained optical flow model to construct multi-frame nonlinear flow-based intermediate frames and confidence maps. These flow-guided frames are then encoded as latent priors to initialize and iteratively guide a pre-trained Video Diffusion model, enabling the diffusion model to preserve dense motion correspondence while improving perceptual realism. To further enhance output quality, we employ confidence maps to fuse structurally reliable flow-based predictions with diffusion-generated details in uncertain regions such as occlusions and object boundaries. Extensive evaluations on challenging benchmarks, including DAVIS, Sintel, and KITTI, demonstrate that SNM-VFI achieves strong perceptual quality, competitive reconstruction accuracy, and robust temporal coherence across diverse motion scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。