通过编辑视频稀疏轨迹实现精准运动修改
MotionV2V: Editing Motion in a Video
- 直接修改输入视频提取的稀疏轨迹来控制运动
- 生成内容相同但运动不同的视频对用于训练
- 支持任意时间点开始的自然运动编辑,用户偏好超65%
尽管生成式视频模型已达到出色的真实感与一致性,将其应用于视频编辑仍面临复杂挑战。现有研究探索了运动可控性以提升文本到视频生成或图像动画能力,但我们指出精确运动控制是极具潜力却未被充分探索的视频编辑范式。本文提出通过直接编辑输入视频中提取的稀疏轨迹来修改视频运动。我们将输入与输出轨迹之间的差异称为‘运动编辑’,并证明该表示结合生成主干网络可实现强大视频编辑能力。为此,我们构建了生成‘运动反事实’的流程,即内容相同但运动不同的视频对,并在该数据集上微调运动条件化的视频扩散架构。我们的方法支持从任意时间戳开始的运动编辑,并实现自然传播。四人头对头用户研究显示,模型获得超过65%的偏好胜率。
原文摘要 · Abstract (English)
While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explored motion controllability as a means to enhance text-to-video generation or image animation; however, we identify precise motion control as a promising yet under-explored paradigm for editing existing videos. In this work, we propose modifying video motion by directly editing sparse trajectories extracted from the input. We term the deviation between input and output trajectories a "motion edit" and demonstrate that this representation, when coupled with a generative backbone, enables powerful video editing capabilities. To achieve this, we introduce a pipeline for generating "motion counterfactuals", video pairs that share identical content but distinct motion, and we fine-tune a motion-conditioned video diffusion architecture on this dataset. Our approach allows for edits that start at any timestamp and propagate naturally. In a four-way head-to-head user study, our model achieves over 65 percent preference against prior work. Please see our project page: https://ryanndagreat.github.io/MotionV2V
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。