让文本生成视频能精准组合多种动作,突破传统方法的运动控制瓶颈。
CoMo: Compositional Motion Customization for Text-to-Video Generation
- 分两阶段学习:先解耦运动与外观,再通过空间隔离合成多动作。
- 在多个数据集上实现最高水平的多动作生成精度,显著优于现有方法。
- 适合需要精细动作控制的视频生成研究者或创意应用开发者。
尽管近期文本到视频模型在生成多样化场景方面表现优异,但在复杂多主体动作的精确控制上仍存在困难。现有单动作定制方法在组合场景中因运动-外观纠缠和多动作融合无效而失效。本文提出CoMo框架,实现文本到视频生成中的组合运动定制,支持单视频内合成多种独立动作。该方法采用两阶段策略:第一阶段通过静态-动态解耦微调,将运动从外观中分离,学习专用运动模块;第二阶段采用即插即用的分-合策略,在去噪过程中通过空间隔离实现动作合成,无需额外训练。为推动该领域研究,我们还构建了新基准与评估指标,用于衡量多动作保真度与融合质量。大量实验表明,CoMo达到当前最优性能,显著提升可控视频生成能力。
原文摘要 · Abstract (English)
While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to address this gap, they fail in compositional scenarios due to two primary challenges: motion-appearance entanglement and ineffective multi-motion blending. This paper introduces CoMo, a novel framework for $\textbf{compositional motion customization}$ in text-to-video generation, enabling the synthesis of multiple, distinct motions within a single video. CoMo addresses these issues through a two-phase approach. First, in the single-motion learning phase, a static-dynamic decoupled tuning paradigm disentangles motion from appearance to learn a motion-specific module. Second, in the multi-motion composition phase, a plug-and-play divide-and-merge strategy composes these learned motions without additional training by spatially isolating their influence during the denoising process. To facilitate research in this new domain, we also introduce a new benchmark and a novel evaluation metric designed to assess multi-motion fidelity and blending. Extensive experiments demonstrate that CoMo achieves state-of-the-art performance, significantly advancing the capabilities of controllable video generation. Our project page is at https://como6.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。