用动态遮罩控制视频生成,低成本实现运动一致性
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
- 通过遮罩运动序列引导生成过程,精准控制前景物体运动轨迹
- 仅需少量训练数据即可实现高质量长视频生成,保持前后一致
- 适合需要可控运动的视频编辑与艺术创作场景
扩散模型为视觉内容创作带来新活力。然而,当前文本到视频生成模型仍面临训练成本高、数据需求大以及文本与前景物体运动不一致等挑战。为此,我们提出基于遮罩引导的视频生成方法,通过遮罩运动序列控制生成过程,仅需有限训练数据。该模型在现有架构基础上引入前景遮罩,实现文本位置精准匹配与运动轨迹控制。利用遮罩运动序列,确保生成序列中前景物体的一致性。此外,结合首帧共享策略与自回归扩展方法,实现更稳定、更长时长的视频生成。大量定性与定量实验表明,该方法在视频编辑和艺术视频生成等任务中表现优异,显著优于以往方法,在一致性和质量上均有提升。生成结果详见补充材料。
原文摘要 · Abstract (English)
Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and difficulties in maintaining consistency between given text and motion of the foreground object. To address these challenges, we propose mask-guided video generation, which can control video generation through mask motion sequences, while requiring limited training data. Our model enhances existing architectures by incorporating foreground masks for precise text-position matching and motion trajectory control. Through mask motion sequences, we guide the video generation process to maintain consistent foreground objects throughout the sequence. Additionally, through a first-frame sharing strategy and autoregressive extension approach, we achieve more stable and longer video generation. Extensive qualitative and quantitative experiments demonstrate that this approach excels in various video generation tasks, such as video editing and generating artistic videos, outperforming previous methods in terms of consistency and quality. Our generated results can be viewed in the supplementary materials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。