arXiv:2412.13190cs.CV2024-12被引 15

让视频补帧更灵活,支持轨迹、关键帧等多方式控制。

MotionBridge: Dynamic Video Inbetweening with Flexible Controls

  • 双分支结构提取多种控制信号,避免歧义。
  • 通过渐进式训练策略,实现多模态控制统一学习。
  • 支持轨迹、遮罩、文本等灵活控制,适合创意视频编辑。

通过生成两个图像帧之间的合理且平滑过渡,视频补帧是视频编辑与长视频合成的关键技术。传统方法难以生成复杂的大范围运动;而现有视频生成技术虽能产出高质量结果,却常缺乏对中间帧细节的精细控制,导致与创作者意图不符。我们提出 MotionBridge,一种统一的视频补帧框架,支持轨迹笔画、关键帧、掩码、引导像素和文本等多种灵活控制。为在统一框架中学习多模态控制,我们设计了双生成器以准确提取控制信号,并采用双分支嵌入编码器缓解歧义问题。进一步引入渐进式训练策略,使模型能逐步学习不同控制模式。大量定性与定量实验表明,该多模态控制机制显著提升了生成视频的动态性、可定制性和上下文一致性。

原文摘要 · Abstract (English)

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While recent video generation techniques are powerful in creating high-quality results, they often lack fine control over the details of intermediate frames, which can lead to results that do not align with the creative mind. We introduce MotionBridge, a unified video inbetweening framework that allows flexible controls, including trajectory strokes, keyframes, masks, guide pixels, and text. However, learning such multi-modal controls in a unified framework is a challenging task. We thus design two generators to extract the control signal faithfully and encode feature through dual-branch embedders to resolve ambiguities. We further introduce a curriculum training strategy to smoothly learn various controls. Extensive qualitative and quantitative experiments have demonstrated that such multi-modal controls enable a more dynamic, customizable, and contextually accurate visual narrative.

视频生成可控生成多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。