世界模型需从生成画面转向可行动模拟,强调物理约束与因果结构。
From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models
- 将世界模型重构为可行动模拟器,注重因果结构与约束
- 在医疗决策中验证,真实价值在于反事实推理而非画面逼真度
- 提出4D结构接口与闭环评估,适合高风险决策场景
世界模型是通过想象未来来实现规划的AI系统,而非依赖即时感知。当前模型存在视觉混淆:高保真视频生成并不等于理解物理与因果动态。我们发现,现代模型虽能精准预测像素,却常违反不变约束、干预下失效,在安全关键决策中崩溃。本综述指出,视觉真实无法可靠代表世界理解。真正有效的世界模型必须编码因果结构、遵守领域特定约束,并在长时程中保持稳定。我们主张将世界模型重定义为可行动模拟器,强调结构化4D接口、约束感知动力学和闭环评估。以医疗决策为认知压力测试——试错不可行且错误不可逆——我们证明,世界模型的价值不在于播放效果是否逼真,而在于支持反事实推理、干预规划与稳健长时前瞻的能力。
原文摘要 · Abstract (English)
A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world models, however, suffer from visual conflation: the mistaken assumption that high-fidelity video generation implies an understanding of physical and causal dynamics. We show that while modern models excel at predicting pixels, they frequently violate invariant constraints, fail under intervention, and break down in safety-critical decision-making. This survey argues that visual realism is an unreliable proxy for world understanding. Instead, effective world models must encode causal structure, respect domain-specific constraints, and remain stable over long horizons. We propose a reframing of world models as actionable simulators rather than visual engines, emphasizing structured 4D interfaces, constraint-aware dynamics, and closed-loop evaluation. Using medical decision-making as an epistemic stress test, where trial-and-error is impossible and errors are irreversible, we demonstrate that a world model's value is determined not by how realistic its rollouts appear, but by its ability to support counterfactual reasoning, intervention planning, and robust long-horizon foresight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。