用遮罩序列精准控制动作,让视频生成更流畅复杂。
Motion Control for Enhanced Complex Action Video Generation
- 引入遮罩序列作为动作条件,弥补文本描述不足
- 支持独立或联合调整文本与动作条件,生成更动态视频
- 可编辑动作条件,适合复杂动作视频创作
现有文本到视频(T2V)模型在生成具有明显或复杂动作的视频时表现不佳,主要受限于文本提示难以精确传达复杂运动细节。为此,我们提出MVideo框架,旨在生成长时长、动作精准流畅的视频。MVideo通过引入遮罩序列作为额外的动作条件输入,克服文本提示的局限性,提供更清晰准确的动作表达。利用GroundingDINO和SAM2等基础视觉模型,自动构建遮罩序列,提升效率与鲁棒性。实验表明,训练后MVideo能有效对齐文本提示与动作条件,生成同时满足二者要求的视频。该双控机制支持独立或联合调整文本与动作条件,实现更动态的视频生成。此外,支持动作条件编辑与组合,促进更复杂动作的视频生成。MVideo显著提升T2V动作生成能力,为当前视频扩散模型中的动作表现树立新基准。项目页面:https://mvideo-v1.github.io/
原文摘要 · Abstract (English)
Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this, we propose a novel framework, MVideo, designed to produce long-duration videos with precise, fluid actions. MVideo overcomes the limitations of text prompts by incorporating mask sequences as an additional motion condition input, providing a clearer, more accurate representation of intended actions. Leveraging foundational vision models such as GroundingDINO and SAM2, MVideo automatically generates mask sequences, enhancing both efficiency and robustness. Our results demonstrate that, after training, MVideo effectively aligns text prompts with motion conditions to produce videos that simultaneously meet both criteria. This dual control mechanism allows for more dynamic video generation by enabling alterations to either the text prompt or motion condition independently, or both in tandem. Furthermore, MVideo supports motion condition editing and composition, facilitating the generation of videos with more complex actions. MVideo thus advances T2V motion generation, setting a strong benchmark for improved action depiction in current video diffusion models. Our project page is available at https://mvideo-v1.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。