让视频生成精准控制运动方向和强度,无需额外训练。
Mojito: Motion Trajectory and Intensity Control for Video Generation
- 用交叉注意力实现无训练的方向引导,高效控制物体运动轨迹。
- 通过光流图调节运动强度,生成符合指定动态的视频。
- 适合需要精细动作控制的视频创作与动画生成场景。
扩散模型在高质量视频生成方面取得显著进展,但如何高效训练具备方向引导和可调运动强度能力的视频扩散模型仍具挑战且研究不足。为此,本文提出Mojito,一种支持运动轨迹与强度控制的文本到视频生成扩散模型。Mojito包含方向运动控制(DMC)模块,利用交叉注意力机制在不需训练的情况下高效引导生成对象的运动方向;同时引入运动强度调制(MIM)模块,通过视频生成的光流图指导不同层级的运动强度。大量实验表明,Mojito在保持高计算效率的同时,能精确实现轨迹与强度控制,生成的运动模式高度匹配预设方向与强度,展现出与真实世界自然运动一致的逼真动态。
原文摘要 · Abstract (English)
Recent advancements in diffusion models have shown great promise in producing high-quality video content. However, efficiently training video diffusion models capable of integrating directional guidance and controllable motion intensity remains a challenging and under-explored area. To tackle these challenges, this paper introduces Mojito, a diffusion model that incorporates both motion trajectory and intensity control for text-to-video generation. Specifically, Mojito features a Directional Motion Control (DMC) module that leverages cross-attention to efficiently direct the generated object's motion without training, alongside a Motion Intensity Modulator (MIM) that uses optical flow maps generated from videos to guide varying levels of motion intensity. Extensive experiments demonstrate Mojito's effectiveness in achieving precise trajectory and intensity control with high computational efficiency, generating motion patterns that closely match specified directions and intensities, providing realistic dynamics that align well with natural motion in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。