arXiv:2505.14167cs.CV2025-05被引 3

让视频生成更懂动作,用参考视频控制运动轨迹

LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer

  • 通过解耦前景与背景,精准提取参考视频的运动信息
  • 在图文与图视频生成中实现零样本运动控制,性能领先
  • 适合需要精确控制动作的视频生成应用

近年来,大规模预训练扩散变换器(DiT)在视频生成领域取得显著进展。尽管现有模型能生成高分辨率、高帧率且多样性强的视频,但在内容细节控制方面仍显不足,尤其难以仅通过文本提示精确控制复杂运动。现有方法在图到视频生成中也缺乏运动控制能力,因参考图像与参考视频中的主体在初始位置、大小和形状上存在差异。为此,我们提出零样本视频生成的运动先验利用框架(LMP)。该框架借助预训练扩散变换器的强大生成能力,使生成视频的运动可参考用户提供的运动视频,涵盖文本到视频与图像到视频生成。我们设计了前景-背景解耦模块以分离运动主体与背景,避免干扰;引入重加权运动迁移模块,实现运动信息的有效传递;并提出外观分离模块,抑制参考主体外观在目标视频中的残留。我们在DAVIS数据集上构建了带详细提示的标注集,并设计评估指标验证方法有效性。大量实验表明,本方法在生成质量、提示-视频一致性及控制能力上均达到当前最佳水平。

原文摘要 · Abstract (English)

In recent years, large-scale pre-trained diffusion transformer models have made significant progress in video generation. While current DiT models can produce high-definition, high-frame-rate, and highly diverse videos, there is a lack of fine-grained control over the video content. Controlling the motion of subjects in videos using only prompts is challenging, especially when it comes to describing complex movements. Further, existing methods fail to control the motion in image-to-video generation, as the subject in the reference image often differs from the subject in the reference video in terms of initial position, size, and shape. To address this, we propose the Leveraging Motion Prior (LMP) framework for zero-shot video generation. Our framework harnesses the powerful generative capabilities of pre-trained diffusion transformers to enable motion in the generated videos to reference user-provided motion videos in both text-to-video and image-to-video generation. To this end, we first introduce a foreground-background disentangle module to distinguish between moving subjects and backgrounds in the reference video, preventing interference in the target video generation. A reweighted motion transfer module is designed to allow the target video to reference the motion from the reference video. To avoid interference from the subject in the reference video, we propose an appearance separation module to suppress the appearance of the reference subject in the target video. We annotate the DAVIS dataset with detailed prompts for our experiments and design evaluation metrics to validate the effectiveness of our method. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in generation quality, prompt-video consistency, and control capability. Our homepage is available at https://vpx-ecnu.github.io/LMP-Website/

视频生成扩散模型运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。