测试视频模型在物理事件中的反事实一致性,发现视角变化严重影响生成质量。
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models

- 通过控制视角、场景等四因素,评估模型对物理事件的响应能力。
- 同一物理事件下,外观和视角变化导致生成质量明显下降。
- 适合关注视频生成鲁棒性与因果推理的研究者使用。
视频预测被视为通向通用世界模型的重要路径,但当前系统是否学习了底层因果结构仍不明确,可能仅依赖表面视觉相关性进行未来预测。我们提出CRONOS,一个基于干预的基准,用于评估反事实物理一致性:即模型对物理事件的预测是否能正确响应视觉输入的可控变化,如场景上下文、视角、物体外观和类别。该基准在逼真的Unreal Engine环境中构建,可生成多样场景与动态的高质量视频。与以往基准不同,CRONOS系统地干预四个关键因素——视角、场景、物体类别和外观,同时保持底层物理事件类型(如碰撞、遮挡或坠落)不变。对近期开源视频生成模型的评估显示,其反事实物理一致性存在显著缺陷:相同物理事件的生成质量受外观、环境及尤其是视角变化影响明显。CRONOS提供了一个可控且可复现的测试平台,用于诊断不同干预下的生成质量变化,为开发跨多种条件保持一致性的模型设立了具体目标。数据集与代码已公开。
原文摘要 · Abstract (English)
Video prediction is increasingly viewed as a path toward generalizable world models, yet it remains unclear whether these systems learn underlying causal structure or merely exploit superficial visual correlations for future prediction. We introduce CRONOS, an intervention-based benchmark designed to evaluate counterfactual physical consistency: whether a model's predictions of physical events respond appropriately to controlled changes in the visual input, such as variations of scene context, viewpoint, object appearance, and object category. Built in a photorealistic Unreal Engine environment, CRONOS enables controlled, high-fidelity generation of videos across diverse scenes and dynamics. In contrast to previous benchmarks, CRONOS systematically intervenes on four key factors - viewpoint, scene, object category, and object appearance - while keeping the underlying physical event type, such as a collision, occlusion, or fall, fixed. Our evaluation of recent open-source video generators reveals substantial failures in counterfactual physical consistency: prediction quality for the same physical event type is affected by appearance, environment, and, particularly by viewpoint changes. CRONOS provides a controlled and reproducible testbed for diagnosing how the quality of generated videos changes for different interventions, establishing a concrete target for developing models that perform consistently across changes of multiple conditions. The dataset and code are available at our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。