arXiv:2509.22642cs.ROcs.CV2025-09被引 59

用200万次机器人交互训练大模型,让AI真正理解物理因果。

WoW: Towards a World omniscient World model Through Embodied Interaction

  • 基于真实世界交互数据训练140亿参数世界模型
  • 模型能预测物理事件概率分布,但需约束避免幻觉
  • 适合研究具身智能与物理推理的开发者

人类通过与世界主动互动获得直观物理认知,这与当前依赖被动观察的视频模型(如Sora)形成鲜明对比,后者难以掌握物理因果关系。我们提出核心假设:真正的物理直觉必须建立在大量、因果丰富的现实世界交互之上。为此,我们构建了WoW,一个基于200万条机器人交互轨迹训练的140亿参数生成式世界模型。实验发现,该模型对物理的理解表现为合理结果的概率分布,导致随机不稳定性与物理幻觉。进一步证明,通过SOPHIA框架——由视觉语言模型代理评估DiT生成输出并迭代优化语言指令,可主动约束模型以提升物理真实性。同时,联合训练的逆动力学模型将优化后的计划转化为可执行机器人动作,实现从想象到行动的闭环。我们建立了WoWBench新基准,聚焦视频生成中的物理一致性和因果推理,结果显示WoW在人类与自动评估中均达领先水平,展现出强大的物理因果、碰撞动力学与物体恒常性能力。本工作系统证实大规模真实交互是培养AI物理直觉的关键。模型、数据与基准将开源。

原文摘要 · Abstract (English)

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore struggle with grasping physical causality. This observation leads to our central hypothesis: authentic physical intuition of the world model must be grounded in extensive, causally rich interactions with the real world. To test this hypothesis, we present WoW, a 14-billion-parameter generative world model trained on 2 million robot interaction trajectories. Our findings reveal that the model's understanding of physics is a probabilistic distribution of plausible outcomes, leading to stochastic instabilities and physical hallucinations. Furthermore, we demonstrate that this emergent capability can be actively constrained toward physical realism by SOPHIA, where vision-language model agents evaluate the DiT-generated output and guide its refinement by iteratively evolving the language instructions. In addition, a co-trained Inverse Dynamics Model translates these refined plans into executable robotic actions, thus closing the imagination-to-action loop. We establish WoWBench, a new benchmark focused on physical consistency and causal reasoning in video, where WoW achieves state-of-the-art performance in both human and autonomous evaluation, demonstrating strong ability in physical causality, collision dynamics, and object permanence. Our work provides systematic evidence that large-scale, real-world interaction is a cornerstone for developing physical intuition in AI. Models, data, and benchmarks will be open-sourced.

具身智能物理推理世界模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。