arXiv:2604.01700cs.CVcs.MM2026-04

让视频插帧模型能自检运动一致性,避免抖动和方向错乱。

Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation

  • 双向循环一致性约束,正反生成路径对称优化
  • 37帧与73帧任务上均超越现有方法,无额外计算开销
  • 适合需要稳定长序列生成的视频修复与动画应用

视频帧插值旨在合成给定端点间的逼真中间帧,同时遵循特定运动语义。尽管生成模型提升了视觉质量,但多数采用单向生成,缺乏自验证时间一致性的机制,常导致运动漂移、方向模糊和边界错位,尤其在长序列中更为明显。受自监督学习中时间循环一致性的启发,我们提出一种新颖的双向框架,强制正向与反向生成轨迹保持对称性。该方法引入可学习的方向标记,显式控制共享主干网络的时序方向,使模型在统一架构中联合优化正向合成与反向重建。这种循环一致性监督作为强正则项,确保生成运动路径逻辑可逆。此外,采用从短到长的课程学习策略,稳定不同持续时间下的动态行为。关键在于,循环约束仅用于训练;推理只需单次前向传播,保持基线模型高效性。大量实验表明,本方法在37帧与73帧任务上均达到当前最优性能,图像质量、运动平滑度与动态控制表现优异,且无额外计算开销。

原文摘要 · Abstract (English)

Video frame interpolation aims to synthesize realistic intermediate frames between given endpoints while adhering to specific motion semantics. While recent generative models have improved visual fidelity, they predominantly operate in a unidirectional manner, lacking mechanisms to self-verify temporal consistency. This often leads to motion drift, directional ambiguity, and boundary misalignment, especially in long-range sequences. Inspired by the principle of temporal cycle-consistency in self-supervised learning, we propose a novel bidirectional framework that enforces symmetry between forward and backward generation trajectories. Our approach introduces learnable directional tokens to explicitly condition a shared backbone on temporal orientation, enabling the model to jointly optimize forward synthesis and backward reconstruction within a single unified architecture. This cycle-consistent supervision acts as a powerful regularizer, ensuring that generated motion paths are logically reversible. Furthermore, we employ a curriculum learning strategy that progressively trains the model from short to long sequences, stabilizing dynamics across varying durations. Crucially, our cyclic constraints are applied only during training; inference requires a single forward pass, maintaining the high efficiency of the base model. Extensive experiments show that our method achieves state-of-the-art performance in imaging quality, motion smoothness, and dynamic control on both 37-frame and 73-frame tasks, outperforming strong baselines while incurring no additional computational overhead.

视频插帧扩散模型循环一致性运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。