无需训练即可精准控制视频运动,支持像素级外观调节。
Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising
- 用剪切拖拽等简单操作生成粗略动画作运动参考
- 双时钟去噪实现关键区域强对齐,整体保持自然动态
- 插件式设计兼容任意主干模型,适合需要精细控制的用户
基于扩散模型的视频生成虽能生成逼真视频,但现有图像或文本条件难以实现精确运动控制。以往运动条件生成方法通常需针对特定模型微调,计算成本高且限制多。本文提出Time-to-Move(TTM),一种无需训练、即插即用的图像到视频(I2V)扩散模型运动与外观控制框架。核心思路是利用用户友好的操作(如剪切拖拽或基于深度的重投影)生成粗略参考动画,将其作为粗粒度运动提示。受SDEdit使用粗略布局进行图像编辑的启发,我们将此方法拓展至视频领域。通过图像条件保留外观,并引入双时钟去噪——一种区域依赖策略,在指定运动区域强制强对齐,其他区域保持灵活性,从而在忠实于用户意图与保持自然动态间取得平衡。该轻量级采样过程修改不增加额外训练或运行开销,兼容任意主干模型。大量实验表明,TTM在物体运动和相机运动基准上达到或超过现有训练型基线的逼真度与运动控制能力。此外,TTM首次实现通过像素级条件进行精确外观控制,突破纯文本提示的局限。项目主页提供视频示例与代码:https://time-to-move.github.io/。
原文摘要 · Abstract (English)
Diffusion-based video generation can create realistic videos, yet existing image- and text-based conditioning fails to offer precise motion control. Prior methods for motion-conditioned synthesis typically require model-specific fine-tuning, which is computationally expensive and restrictive. We introduce Time-to-Move (TTM), a training-free, plug-and-play framework for motion- and appearance-controlled video generation with image-to-video (I2V) diffusion models. Our key insight is to use crude reference animations obtained through user-friendly manipulations such as cut-and-drag or depth-based reprojection. Motivated by SDEdit's use of coarse layout cues for image editing, we treat the crude animations as coarse motion cues and adapt the mechanism to the video domain. We preserve appearance with image conditioning and introduce dual-clock denoising, a region-dependent strategy that enforces strong alignment in motion-specified regions while allowing flexibility elsewhere, balancing fidelity to user intent with natural dynamics. This lightweight modification of the sampling process incurs no additional training or runtime cost and is compatible with any backbone. Extensive experiments on object and camera motion benchmarks show that TTM matches or exceeds existing training-based baselines in realism and motion control. Beyond this, TTM introduces a unique capability: precise appearance control through pixel-level conditioning, exceeding the limits of text-only prompting. Visit our project page for video examples and code: https://time-to-move.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。