无需训练即可精准控制多事件视频生成过程。
TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation

- 通过分段掩码与跨事件提示融合,逐步引导生成
- 在8项指标上达当前最优,支持一致性与事件分离权衡
- 适合需要灵活控制多事件视频生成的研究者
文本到视频生成在长时序、多事件场景下面临挑战。受扩散过程内在机制启发,我们研究了视频扩散变压器(DiT)的去噪轨迹,发现条件文本影响生成的关键转折点:从全局布局到细粒度细节。基于此,提出TunerDiT——一种无需额外训练的渐进式控制方法。其包含两个控制机制:(1) 事件分段掩码,强制事件边界同时保留跨事件过渡区域;(2) 跨事件提示融合,注入邻近事件语义用于后期优化。我们构建了一个自标注提示数据集Meve用于多事件生成评测。TunerDiT在8项指标上达到当前最优,可调节视频一致性与事件分离性之间的平衡。随着事件数量增加,文本对齐性能提升,显示出良好的扩展潜力。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation faces challenging questions when generating videos with long horizons containing multiple events. Inspired by the intrinsics of the diffusion process, we probe video diffusion transformers (DiTs) and uncover intrinsic turning points in the DiT denoising trajectory where conditioning text affects generation from global layout to fine-grained details. Building on this finding, we present TunerDiT, a simple yet effective progressive steering method that requires no additional training for multi-event generation. TunerDiT comprises two steering handles: (1) Event-Partitioned Masking that enforces event boundaries while allowing cross-event transition bands; (2) Cross-Event Prompt Fusion that injects neighboring event semantics for late-stage refinement. We contribute a self-curated prompt suite for benchmarking multi-event generation, i.e., Meve. TunerDiT achieves state-of-the-art performance across 8 metrics and offers a tunable trade-off between video consistency and event separation, compared with other training-free methods. The improvement in text alignment increases with the event count, indicating a scaling possibility with increasing event count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。