探究多事件视频生成中事件切换的时间与位置规律
When and Where do Events Switch in Multi-Event Video Generation?
- 构建自标注提示集MEve,系统评估事件切换时机
- 发现去噪步骤早期干预和分块层调控是关键
- 为未来可控多事件生成模型提供设计依据
文本到视频(T2V)生成在应对长视频中多个连续事件的时序一致性与内容可控性挑战时迅速发展。现有方法在扩展至多事件生成时,忽略了事件切换的内在机制。本文旨在回答核心问题:多事件提示如何在T2V生成过程中控制事件转换的时间与位置。提出MEve——一个自标注的提示数据集,用于评估多事件文本到视频生成,并对两类代表性模型(OpenSora与CogVideoX)进行系统研究。大量实验表明,去噪过程中的早期干预以及分块式模型层的作用至关重要,揭示了多事件视频生成的核心因素,也为未来模型的多事件条件控制提供了可能方向。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and controllable content. Existing methods that extend to multi-event generation omit an inspection of the intrinsic factor in event shifting. The paper aims to answer the central question: When and where multi-event prompts control event transition during T2V generation. This work introduces MEve, a self-curated prompt suite for evaluating multi-event text-to-video (T2V) generation, and conducts a systematic study of two representative model families, i.e., OpenSora and CogVideoX. Extensive experiments demonstrate the importance of early intervention in denoising steps and block-wise model layers, revealing the essential factor for multi-event video generation and highlighting the possibilities for multi-event conditioning in future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。