无需训练即可生成多事件视频,让每个镜头精准对应提示词
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
- 通过事件对齐查询控制,让注意力聚焦于相关事件
- 自适应调节强度,保持时间连贯性和画面质量
- 适合需要精准叙事的视频生成场景
近期文本到视频的扩散模型已实现高质量、时序连贯的视频生成。然而,现有模型主要针对单事件生成优化,处理多事件提示时缺乏时间锚定,常导致场景混合或坍缩,破坏叙事逻辑。为此,我们提出 SwitchCraft,一种无需训练的多事件视频生成框架。核心思路是:时间上统一注入提示词会忽略事件与帧的对应关系。为此,我们引入事件对齐查询控制(EAQS),引导帧级注意力对齐相关事件提示;同时提出自适应平衡强度求解器(ABSS),动态调节控制强度以兼顾时间一致性与视觉保真度。大量实验表明,相比基线方法,SwitchCraft显著提升提示对齐度、事件清晰度和场景连贯性,为多事件视频生成提供简单而有效的解决方案。
原文摘要 · Abstract (English)
Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for single-event generation. When handling multi-event prompts, without explicit temporal grounding, such models often produce blended or collapsed scenes that break the intended narrative. To address this limitation, we present SwitchCraft, a training-free framework for multi-event video generation. Our key insight is that uniform prompt injection across time ignores the correspondence between events and frames. To this end, we introduce Event-Aligned Query Steering (EAQS), which steers frame-level attention to align with relevant event prompts. Furthermore, we propose Auto-Balance Strength Solver (ABSS), which adaptively balances steering strength to preserve temporal consistency and visual fidelity. Extensive experiments demonstrate that SwitchCraft substantially improves prompt alignment, event clarity, and scene consistency compared with existing baselines, offering a simple yet effective solution for multi-event video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。