arXiv:2412.18597cs.CVcs.AI2024-12CVPR被引 80

无需训练即可生成多提示连贯长视频,解决跨提示过渡生硬问题。

DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

论文配图:DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
图 1 · 摘自论文原文
  • 利用注意力机制实现多提示间语义精准控制,共享注意力提升一致性。
  • 生成视频在多提示场景下过渡自然、物体运动连贯,性能领先现有方法。
  • 提出新基准MPVBench,专用于评估多提示视频生成效果,适合视频编辑研究者。

Sora类视频生成模型采用多模态扩散变换器(MM-DiT)架构取得显著进展。然而,当前模型多聚焦单提示生成,难以处理多个连续提示的复杂动态场景,导致生成内容不连贯。尽管已有研究尝试多提示生成,仍面临严格训练数据依赖、提示跟随能力弱及过渡不自然等问题。为此,本文首次提出无需训练的多提示视频生成方法DiTCtrl,将任务视为具有平滑过渡的时序视频编辑。通过分析MM-DiT的注意力机制,发现3D全注意力与UNet类扩散模型中的交叉/自注意力块行为相似,可实现跨提示的掩码引导语义控制,并共享注意力以支持多提示生成。实验表明,该方法在无额外训练条件下实现流畅过渡与一致物体运动,且在新构建的多提示视频生成基准MPVBench上达到先进水平。

原文摘要 · Abstract (English)

Sora-like video generation models have achieved remarkable progress with a Multi-Modal Diffusion Transformer MM-DiT architecture. However, the current video generation models predominantly focus on single-prompt, struggling to generate coherent scenes with multiple sequential prompts that better reflect real-world dynamic scenarios. While some pioneering works have explored multi-prompt video generation, they face significant challenges including strict training data requirements, weak prompt following, and unnatural transitions. To address these problems, we propose DiTCtrl, a training-free multi-prompt video generation method under MM-DiT architectures for the first time. Our key idea is to take the multi-prompt video generation task as temporal video editing with smooth transitions. To achieve this goal, we first analyze MM-DiT's attention mechanism, finding that the 3D full attention behaves similarly to that of the cross/self-attention blocks in the UNet-like diffusion models, enabling mask-guided precise semantic control across different prompts with attention sharing for multi-prompt video generation. Based on our careful design, the video generated by DiTCtrl achieves smooth transitions and consistent object motion given multiple sequential prompts without additional training. Besides, we also present MPVBench, a new benchmark specially designed for multi-prompt video generation to evaluate the performance of multi-prompt generation. Extensive experiments demonstrate that our method achieves state-of-the-art performance without additional training.

视频生成扩散模型多提示注意力控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。