测试视频生成模型对世界状态演化的推理能力,发现视觉真实不等于逻辑合理。
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

- 将视频生成任务转化为世界状态预测,评估模型对动作后世界变化的合理性。
- 436个测试用例覆盖4个维度22个子类,揭示主流模型在因果与信息一致性上普遍失败。
- 引入人类对齐的多维评估体系,适合研究具身认知与可信生成的学者使用。
商业视频生成系统如Seedance2.0和Veo3.1迅速发展,促使人们认为视频生成器正演变为“世界模拟器”。但学界仍缺乏直接检验模型能否推理世界状态随时间演变的基准。我们提出WorldReasonBench,将视频生成评估重构为世界状态预测:给定初始状态和动作,模型能否生成在物理、社会、逻辑和信息层面均一致的未来视频?该基准包含436个精心设计的测试用例,涵盖四个推理维度与22个子类别,并配有结构化标注的问答数据。我们采用人类对齐的两阶段评估方法:过程感知推理验证通过结构化问答与推理阶段诊断检测时间与因果错误;多维质量评估则对推理质量、时间一致性与视觉美感进行评分,支持排名与奖励建模。此外,我们构建了WorldRewardBench,一个包含约6,000组专家标注配对的偏好基准,覆盖1,400余段视频,支持成对与单点奖励模型评估。在现代视频生成模型中,结果揭示了视觉逼真性与世界推理能力之间的持续鸿沟:视频可能看似合理,但在动态、因果或信息保留上仍存在缺陷。我们将开放基准与评估工具包,以支持社区开展真正具备世界意识的视频生成研究。
原文摘要 · Abstract (English)
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。