测试视频模型在物理因果场景下的推理能力,发现现有模型普遍表现不佳。
What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

- 设计319组对比提示,仅改变一个物理变量,检验视频输出是否符合物理规律。
- 九个顶尖模型平均配对得分不足52%,开源模型仅28%左右。
- 模型表现更依赖视觉明显性而非物理复杂度,细微变化常被忽略。
视频生成模型越来越多地被用作驾驶和机器人操作等具身任务的世界模拟器。关键不在于单个视频是否看起来真实,而在于输入变化时输出是否按物理规律相应改变。为此,我们给模型提供两组描述相同场景但仅一个物理细节不同的提示,检查生成的视频是否如物理预测般产生差异。提示间措辞差异极小,但正确物理差异不明显。若模型未能捕捉这一差异,仍可能生成各自看似合理的视频,而现有基准逐帧评分无法发现此类失败。我们提出What-If World,基于nuScenes和DROID的真实图像构建319组提示对,按驾驶与操作共有的六类物理变量分类组织。每对视频采用APEO四部分评分标准:是否遵循提示(Adherence)、是否物理一致(Physics)、是否保持共享场景(Environment)、是否出现正确差异(Outcome)。在九个前沿模型中,无一超过52%的配对得分,开源模型集中在28%。所有模型在大量因果干预下均表现失败,表明其在支持动作条件模拟或基于模型规划前仍有巨大提升空间。表现较好的案例多与干预的视觉显著性相关,而非其物理可解性;某些视觉细微干预得分低至14.2%,而显著干预最高达40.4%。
原文摘要 · Abstract (English)
Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video looks right, but whether the model's output changes when its input changes. We test this by giving a model two prompts describing the same scene with one physical detail varied, and checking whether the two videos diverge the way physics predicts. The wording difference between the prompts is small by design, since only one variable is changed, but the correct physical difference is not. A model that misses this can still produce two videos that each look plausible individually, and existing benchmarks score videos one at a time and cannot detect this failure. We introduce What-If World, 319 such prompt pairs built on real frames from nuScenes and DROID, organized by a taxonomy of six physical variables shared across driving and manipulation. Each pair is scored with APEO, a four-part rubric checking whether each video follows its prompt (Adherence), is physically consistent (Physics), preserves the shared scene (Environment), and ends in the correct difference (Outcome). Across nine state-of-the-art models, no system exceeds 52% on the paired score, and open-source models cluster near 28%. Every model tested fails on a large fraction of causal interventions, indicating substantial room before these models can reliably support action-conditioned simulation or model-based planning. Where models do score well, performance appears to track the visual prominence of the intervention rather than the tractability of its underlying physics. Some visually subtle interventions score as low as 14.2%, while visually pronounced ones reach 40.4%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。