arXiv:2602.12679cs.CV2026-02中稿 · ICLR被引 5

通过蒸馏运动残差提升视频生成的时序连贯性

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening

  • 在采样时将正向路径的运动残差蒸馏到反向路径
  • 生成结果时序连续性显著提升,减少视觉伪影
  • 适合需要高质量中间帧生成的应用场景

图像到视频(I2V)扩散模型的进展推动了生成式插帧技术的发展,旨在生成两个关键帧之间的语义合理帧。目前流行的推理阶段采样策略利用大规模预训练I2V模型的生成先验,无需额外训练。然而,现有方法在并行融合或串行交替正向与反向路径时,常因两条路径间的运动先验错位导致时间不连续和不良视觉伪影。这是因为每条路径仅遵循自身条件帧所诱导的运动先验。本文提出运动先验蒸馏(MPD),一种简单有效的推理阶段蒸馏技术,通过将正向路径的运动残差蒸馏至反向路径,抑制双向不匹配。该方法可主动避免对末端条件路径的去噪,从而消除路径歧义,保留正向运动先验,生成更时序连贯的插帧结果。我们在标准基准上进行了定量评估,并开展广泛用户研究,验证了方法在实际场景中的有效性。

原文摘要 · Abstract (English)

Recent progress in image-to-video (I2V) diffusion models has significantly advanced the field of generative inbetweening, which aims to generate semantically plausible frames between two keyframes. In particular, inference-time sampling strategies, which leverage the generative priors of large-scale pre-trained I2V models without additional training, have become increasingly popular. However, existing inference-time sampling, either fusing forward and backward paths in parallel or alternating them sequentially, often suffers from temporal discontinuities and undesirable visual artifacts due to the misalignment between the two generated paths. This is because each path follows the motion prior induced by its own conditioning frame. In this work, we propose Motion Prior Distillation (MPD), a simple yet effective inference-time distillation technique that suppresses bidirectional mismatch by distilling the motion residual of the forward path into the backward path. Our method can deliberately avoid denoising the end-conditioned path which causes the ambiguity of the path, and yield more temporally coherent inbetweening results with the forward motion prior. We not only perform quantitative evaluations on standard benchmarks, but also conduct extensive user studies to demonstrate the effectiveness of our approach in practical scenarios.

视频生成扩散模型插帧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。