arXiv:2501.07563cs.CV2025-01被引 17

无需训练即可实现精准运动引导的连贯视频生成

Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

  • 基于初始噪声与新型运动一致性损失,实现无额外训练的运动控制
  • 在多种运动任务中显著提升帧间时序一致性
  • 适合希望快速部署、兼容主流模型的视频生成研究者

本文针对运动引导下视频生成的时序一致性挑战提出解决方案。现有训练自由方法虽避免了模型结构调整或额外训练,但仍难以保证帧间连贯性或精确跟随引导运动。本文提出一种简单有效的方法:结合初始噪声策略与新型运动一致性损失,通过捕捉视频扩散模型中间特征的帧间特征相关模式来表征参考视频的运动模式,并设计该损失以在生成视频中保持相似的相关模式,利用其在潜在空间的梯度引导生成过程,实现精准运动控制。大量实验表明,该方法在多种运动控制任务中显著提升了时序一致性,建立了高效连贯视频生成的新标准。

原文摘要 · Abstract (English)

In this paper, we address the challenge of generating temporally consistent videos with motion guidance. While many existing methods depend on additional control modules or inference-time fine-tuning, recent studies suggest that effective motion guidance is achievable without altering the model architecture or requiring extra training. Such approaches offer promising compatibility with various video generation foundation models. However, existing training-free methods often struggle to maintain consistent temporal coherence across frames or to follow guided motion accurately. In this work, we propose a simple yet effective solution that combines an initial-noise-based approach with a novel motion consistency loss, the latter being our key innovation. Specifically, we capture the inter-frame feature correlation patterns of intermediate features from a video diffusion model to represent the motion pattern of the reference video. We then design a motion consistency loss to maintain similar feature correlation patterns in the generated video, using the gradient of this loss in the latent space to guide the generation process for precise motion control. This approach improves temporal consistency across various motion control tasks while preserving the benefits of a training-free setup. Extensive experiments show that our method sets a new standard for efficient, temporally coherent video generation.

视频生成扩散模型运动引导时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。