arXiv:2603.04899cs.CV2026-03

提升慢速视频插帧质量,保持细节与运动一致性

FC-VFI: Faithful and Consistent Video Frame Interpolation for High-FPS Slow Motion Video Generation

  • 通过潜空间时序建模继承起始帧细节,结合语义匹配线引导运动
  • 支持4倍和8倍插帧,30帧升至120/240帧,分辨率2560×1440
  • 适合高质量慢动作生成,尤其对结构一致性要求高的场景

大规模预训练视频扩散模型在视频插帧上表现优异,但依赖内在生成先验,导致高保真度帧生成困难,难以保留起始与结束帧的细节。现有方法多依赖运动控制保证时间一致性,但密集光流易出错,稀疏点缺乏结构上下文。本文提出FC-VFI,实现忠实且一致的视频帧插值,支持4倍和8倍插帧,将30 FPS提升至120和240 FPS,分辨率2560×1440,同时保持视觉保真度与运动一致性。提出潜空间序列时序建模策略,继承起始与结束帧的保真线索;引入语义匹配线实现结构感知运动引导;设计时序差分损失以缓解时间不一致性。大量实验表明,FC-VFI在多种场景下均具备优异性能与结构完整性。

原文摘要 · Abstract (English)

Large pre-trained video diffusion models excel in video frame interpolation but struggle to generate high fidelity frames due to reliance on intrinsic generative priors, limiting detail preservation from start and end frames. Existing methods often depend on motion control for temporal consistency, yet dense optical flow is error-prone, and sparse points lack structural context. In this paper, we propose FC-VFI for faithful and consistent video frame interpolation, supporting \(4\times\)x and \(8\times\) interpolation, boosting frame rates from 30 FPS to 120 and 240 FPS at \(2560\times 1440\)resolution while preserving visual fidelity and motion consistency. We introduce a temporal modeling strategy on the latent sequences to inherit fidelity cues from start and end frames and leverage semantic matching lines for structure-aware motion guidance, improving motion consistency. Furthermore, we propose a temporal difference loss to mitigate temporal inconsistencies. Extensive experiments show FC-VFI achieves high performance and structural integrity across diverse scenarios.

视频插帧扩散模型慢动作生成运动一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。