让AI通过交互学习物理规律,像人一样越练越强。
IPR-1: Interactive Physical Reasoner
- 用世界模型评估并强化视觉语言模型的决策策略。
- 在1000+种游戏上训练,零样本迁移至未见过的游戏。
- 强调物理因果关系,适合研究具身智能与通用推理。
人类通过观察、互动环境并内化物理和因果关系来学习。本文探讨智能体是否也能通过交互获得类人推理能力,并随经验持续提升。为此,我们构建了包含1000+种异构游戏的G2U基准,存在显著视觉域差异。现有方法(如VLMs、世界模型)因不聚焦核心机制而过度依赖视觉细节,难以捕捉底层物理与因果关系:VLM/VLA缺乏交互中的前瞻推理,世界模型则模仿视觉模式而非分析物理本质。为此,我们提出IPR(Interactive Physical Reasoner),利用世界模型滚动生成的评分来优化VLM策略,并引入PhysCode——一种将语义意图与动态行为对齐的物理导向动作编码,实现预测与推理的统一动作空间。在1000+个游戏上预训练后,IPR在从直觉到目标驱动的各类任务中表现稳健,甚至整体超越GPT-5。性能随训练游戏数和交互步数增加而提升,且具备零样本迁移至未见游戏的能力。结果表明,以物理为核心进行交互是实现持续物理推理的有效路径。更多演示与项目详情见 https://mybearyzhang.github.io/ipr-1。
原文摘要 · Abstract (English)
Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarly acquire human-like reasoning from interaction and keep improving with more experience. To study this, we introduce a Game-to-Unseen (G2U) benchmark of 1,000+ heterogeneous games that exhibit significant visual domain gaps. Existing approaches, including VLMs and world models, struggle to capture underlying physics and causality since they are not focused on core mechanisms and overfit to visual details. VLM/VLA agents reason but lack look-ahead in interactive settings, while world models imagine but imitate visual patterns rather than analyze physics and causality. We therefore propose IPR (Interactive Physical Reasoner), using world-model rollouts to score and reinforce a VLM's policy, and introduce PhysCode, a physics-centric action code aligning semantic intent with dynamics to provide a shared action space for prediction and reasoning. Pretrained on 1,000+ games, our IPR performs robustly on levels from primitive intuition to goal-driven reasoning, and even surpasses GPT-5 overall. We find that performance improves with more training games and interaction steps, and that the model also zero-shot transfers to unseen games. These results support physics-centric interaction as a path to steadily improving physical reasoning. Further demos and project details can be found at https://mybearyzhang.github.io/ipr-1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。