评测文生视频模型的叙事连贯性,发现现有模型在多动作序列中表现不佳。
SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models
- 构建包含320个复杂叙事提示的数据集,覆盖多种情节类型。
- 设计动态时间图评估方法,与人工标注相关性高。
- 揭示当前模型在物体状态一致性、物理合理性及动作时序上的缺陷。
文生视频(T2V)模型在生成视觉吸引人的视频方面取得了显著进展,但在生成需逻辑推进的多事件连续叙事方面仍存在困难。现有基准主要关注视觉质量,缺乏对长序列叙事连贯性的评估。为此,我们提出SeqBench,一个全面评估T2V生成中叙事连贯性的基准。SeqBench包含320个精心设计的提示,涵盖多种叙事复杂度,共生成2,560个由8个先进T2V模型生成的人工标注视频。我们还设计了一种基于动态时间图(DTG)的自动评估指标,能高效捕捉长程依赖和时序关系,同时保持计算效率。该指标与人工标注具有强相关性。通过系统评估,我们发现当前T2V模型存在关键局限:无法在多动作序列中维持物体状态一致,多物体场景中出现物理不合理结果,难以保持动作间真实的时间与顺序关系。SeqBench提供了首个系统化评估叙事连贯性的框架,并为未来模型提升序列推理能力提供明确方向。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through multiple events. Existing T2V benchmarks primarily focus on visual quality metrics but fail to evaluate narrative coherence over extended sequences. To bridge this gap, we present SeqBench, a comprehensive benchmark for evaluating sequential narrative coherence in T2V generation. SeqBench includes a carefully designed dataset of 320 prompts spanning various narrative complexities, with 2,560 human-annotated videos generated from 8 state-of-the-art T2V models. Additionally, we design a Dynamic Temporal Graphs (DTG)-based automatic evaluation metric, which can efficiently capture long-range dependencies and temporal ordering while maintaining computational efficiency. Our DTG-based metric demonstrates a strong correlation with human annotations. Through systematic evaluation using SeqBench, we reveal critical limitations in current T2V models: failure to maintain consistent object states across multi-action sequences, physically implausible results in multi-object scenarios, and difficulties in preserving realistic timing and ordering relationships between sequential actions. SeqBench provides the first systematic framework for evaluating narrative coherence in T2V generation and offers concrete insights for improving sequential reasoning capabilities in future models. Please refer to https://videobench.github.io/SeqBench.github.io/ for more details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。