测试视频生成模型能否真正理解现实因果关系。
Thinking in Video: Can Video Generators Really Reason About the Real World?

- 设计双视角评估框架,从感知与生成两方面检验因果推理能力。
- 发现开源模型虽能生成合理画面,却几乎无因果理解能力。
- 揭示语音视觉不一致现象:说对了但画不对,质疑其世界模拟真实性。
近期世界模型与视频生成的发展催生了一种新推理范式——'思维在视频中'(Thinking in Video),即以视频为媒介进行因果思考的构建、扩展与验证。然而该潜力尚未被证实:令人信服的生成结果可能源于记忆而非因果理解,现有指标也割裂了感知保真度与语义逻辑。为此,我们提出因果-生成双判别评估框架(CGDJ),从两个角度审计世界模型一致性。显式因果感知通过时空扁平化视觉问答测试模型是否将视频场景视为推理问题;隐式生成感知-预测差距则评估模型能否生成符合因果逻辑的未来视频。对代表性开闭源生成器的应用表明存在显著感知-预测差距:开源模型虽能生成合理动态,但显式因果感知接近零;先进闭源系统虽表现更好,但推理与生成仍不完全对齐。进一步分析揭示音视频错位现象——模型口头表达正确因果逻辑的能力强于视觉呈现,挑战了其作为'世界模拟器'的叙事。
原文摘要 · Abstract (English)
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, while existing metrics separate perceptual fidelity from semantic logic. To evaluate whether video generators support such reasoning, we introduce the Causal-Generative Dual-Judge (CGDJ), auditing World Model Consistency from two perspectives. Explicit Causal Perception tests whether a generator reads a video scenario as a reasoning problem through spatio-temporal flattened visual question answering, while Implicit Generative Perception-Prediction Gap evaluates whether it renders the causal consequence as a consistent future video. Applying CGDJ to representative open- and closed-source generators reveals a clear Perception-Prediction Gap: open-source models produce plausible dynamics despite near-zero explicit causal perception, whereas advanced closed-source systems show stronger but still limited alignment between reasoning and generation. Further analysis exposes audio-visual misalignment, where models verbalize correct causal logic more reliably than they render it, challenging the "world simulator" narrative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。