新基准VideoVerse测试文生视频模型是否具备世界模型能力。
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
- 构建事件级时序因果数据集,评估模型对复杂时间逻辑的理解。
- 覆盖300个提示、815个事件、793个问题,涵盖动态静态特性。
- 适合关注视频生成真实性与世界知识建模的研究者使用。
近期文本到视频(T2V)生成技术快速发展,使模型具备更强的世界模型能力,现有评估基准已难以区分先进T2V模型。首先,当前评价维度如帧级美学质量与时间一致性,已无法有效区分顶尖模型;其次,事件级时间因果性——区别于其他模态的关键特性——仍基本未被探索;第三,现有基准缺乏对世界知识的系统性评估,而这是构建世界模型的核心能力。为此,我们提出VideoVerse,一个聚焦评估当前T2V模型是否具备理解复杂时间因果性与世界知识以生成视频能力的综合性基准。我们收集了跨多个领域的代表性视频,提取其具有内在时间因果性的事件级描述,并由独立标注者改写为文生视频提示。每个提示设计十项评估维度,共形成300个提示、815个事件和793个评估问题。基于现代视觉语言模型,构建人类偏好对齐的问答式评估流程,系统性地评测主流开源与闭源T2V系统,揭示了当前模型与理想世界建模能力之间的差距。
原文摘要 · Abstract (English)
The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making the existing benchmarks increasingly insufficient to evaluate state-of-the-art T2V models. First, current evaluation dimensions, such as per-frame aesthetic quality and temporal consistency, are no longer able to differentiate state-of-the-art T2V models. Second, event-level temporal causality-an essential property that differentiates videos from other modalities-remains largely unexplored. Third, existing benchmarks lack a systematic assessment of world knowledge, which are essential capabilities for building world models. To address these issues, we introduce VideoVerse, a comprehensive benchmark focusing on evaluating whether the current T2V model could understand complex temporal causality and world knowledge to synthesize videos. We collect representative videos across diverse domains and extract their event-level descriptions with inherent temporal causality, which are then rewritten into text-to-video prompts by independent annotators. For each prompt, we design ten evaluation dimensions covering dynamic and static properties, resulting in 300 prompts, 815 events, and 793 evaluation questions. Consequently, a human preference-aligned QA-based evaluation pipeline is developed by using modern vision-language models to systematically benchmark leading open- and closed-source T2V systems, revealing the current gap between T2V models and desired world modeling abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。