多模态控制的视频插帧,实现精细动作与内容同步生成。
MultiCOIN: Multi-Modal COntrollable Video INbetweening
- 用点状表示统一多模态控制信号,支持深度、轨迹、文本等指令
- 分离内容与运动分支生成,提升细节可控性与自然度
- 适合影视编辑、动画创作等需要精准动态控制的场景
视频插帧可生成两帧图像间的流畅自然过渡,是视频编辑与长序列视频合成的重要工具。现有方法难以生成大尺度、复杂或精细的动作,且对用户意图适应性差,缺乏对中间帧细节的精细控制,导致与创作意图不一致。为此,我们提出 MultiCOIN 框架,支持深度变化、分层、运动轨迹、文本提示及目标区域定位等多种多模态控制,兼顾灵活性、易用性与细粒度精度。采用扩散变换器(DiT)作为视频生成模型,因其在生成高质量长视频方面表现优异。为兼容多模态控制,将所有运动控制映射为统一的稀疏点状输入。针对不同粒度的控制,将内容与运动控制分别编码,通过两个生成器分别处理运动与内容特征,引导去噪过程。进一步设计分阶段训练策略,使模型逐步学习多模态控制能力。大量定性与定量实验表明,多模态控制显著提升了视觉叙事的动态性、可定制性与上下文准确性。
原文摘要 · Abstract (English)
Video inbetweening creates smooth and natural transitions between two image frames, making it an indispensable tool for video editing and long-form video synthesis. Existing works in this domain are unable to generate large, complex, or intricate motions. In particular, they cannot accommodate the versatility of user intents and generally lack fine control over the details of intermediate frames, leading to misalignment with the creative mind. To fill these gaps, we introduce MultiCOIN, a video inbetweening framework that allows multi-modal controls, including depth transition and layering, motion trajectories, text prompts, and target regions for movement localization, while achieving a balance between flexibility, ease of use, and precision for fine-grained video interpolation. To achieve this, we adopt the Diffusion Transformer (DiT) architecture as our video generative model, due to its proven capability to generate high-quality long videos. To ensure compatibility between DiT and our multi-modal controls, we map all motion controls into a common sparse and user-friendly point-based representation as the video/noise input. Further, to respect the variety of controls which operate at varying levels of granularity and influence, we separate content controls and motion controls into two branches to encode the required features before guiding the denoising process, resulting in two generators, one for motion and the other for content. Finally, we propose a stage-wise training strategy to ensure that our model learns the multi-modal controls smoothly. Extensive qualitative and quantitative experiments demonstrate that multi-modal controls enable a more dynamic, customizable, and contextually accurate visual narrative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。