arXiv:2502.13234cs.CVcs.AI2025-02被引 3

通过特征匹配实现文本生成视频的精准运动控制

MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching

  • 在特征层面匹配运动信息,避免像素级重建
  • 相比现有方法,运动还原更准确且无内容泄露
  • 适合需要精细动作控制的视频生成场景

文本到视频(T2V)扩散模型在从文本提示生成真实视频方面展现出强大能力。然而,仅依赖文本描述难以精确控制物体运动和镜头构图。本文提出MotionMatcher,一种基于参考视频进行运动定制的框架。不同于现有方法通过微调模型来重建参考视频的帧间差异,我们发现该策略易导致参考视频内容泄露,且难以捕捉复杂运动。为此,MotionMatcher在特征层面微调预训练的T2V扩散模型,比较高层时空运动特征而非像素级差异,确保精确运动学习。为兼顾内存效率与可访问性,我们利用预训练的T2V模型自身蕴含的丰富视频运动先验知识来提取这些特征。实验表明,该框架在运动定制任务上达到当前最优性能,验证了其设计有效性。

原文摘要 · Abstract (English)

Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects movements and camera framing. In this work, we tackle the motion customization problem, where a reference video is provided as motion guidance. While most existing methods choose to fine-tune pre-trained diffusion models to reconstruct the frame differences of the reference video, we observe that such strategy suffer from content leakage from the reference video, and they cannot capture complex motion accurately. To address this issue, we propose MotionMatcher, a motion customization framework that fine-tunes the pre-trained T2V diffusion model at the feature level. Instead of using pixel-level objectives, MotionMatcher compares high-level, spatio-temporal motion features to fine-tune diffusion models, ensuring precise motion learning. For the sake of memory efficiency and accessibility, we utilize a pre-trained T2V diffusion model, which contains considerable prior knowledge about video motion, to compute these motion features. In our experiments, we demonstrate state-of-the-art motion customization performances, validating the design of our framework.

视频生成扩散模型运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。