无身体的GPT-5.1竟能在真实环境中导航抓物,展现物理理解能力。
Embodied GPT-5.1: Evidence of a World Model?

- 用低分辨率第一视角图像和离散动作控制机器人,无仿真训练
- 能记住物体位置、推断自身动作后果,完成碰撞后回看等连贯行为
- 挑战了“身体是智能前提”的传统观点,适合研究具身智能的学者
本探索性研究考察大型多模态语言模型GPT-5.1是否可在无先前具身经验、未在模拟环境训练、无感官运动经验的情况下,作为实体移动机器人的高层控制器。仅使用低分辨率第一人称图像和离散动作集,模型被要求执行导航与目标物体交互任务,如定位并接触目标玩具。在多次试验中,GPT-5.1展现出空间推理与物理理解的涌现能力,包括在物体离开视野后保持其位置的短期记忆、推断自身运动的物理后果,并执行如撞击物体后回退以视觉验证结果的连贯动作序列。同时,模型也暴露出效率不足与感知局限,如对齐策略不精确及对远距离干扰物的误识别。总体表明,尽管缺乏具身训练,GPT-5.1在具身场景中表现出类世界模型的行为特征,这一发现挑战了认知科学与机器人学中关于物理身体是发展此类智能必要前提的长期观点。研究结果激励对大语言模型中物理理解的涌现性、局限性与鲁棒性的深入探究。
原文摘要 · Abstract (English)
This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot despite having no prior embodiment, no training in simulated environments, and no exposure to sensorimotor experience. Using only low-resolution first-person images and a discrete action set, the model was tasked with navigation and object-directed behaviors such as locating and contacting a target toy. Across multiple trials, GPT-5.1 demonstrated emergent capabilities that suggest elements of spatial reasoning and physical understanding. These included maintaining short-term memory of object locations after they left the camera frame, inferring the physical consequences of its own movements, and executing coherent action sequences such as colliding with an object and reversing to visually verify the outcome. At the same time, the model displayed inefficiencies and perceptual limitations, including imprecise alignment strategies and occasional misidentification of distant distractors. Overall, the results indicate that GPT-5.1 exhibits signs of world-model-like behavior in an embodied setting, despite the absence of any embodiment-related training, a finding that challenges long-standing views in cognitive science and robotics which hold that a physical body is a necessary prerequisite for developing such forms of intelligence. The findings motivate deeper investigation into the emergence, limits, and robustness of physical understanding in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。