arXiv:2512.05955cs.ROcs.CV2025-12中稿 · CVPR被引 6

让视觉语言模型通过实时仿真具备物理推理能力,实现精准机器人操作

SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models

  • 用仿真构建物理世界模型,让VLM在推理时考虑真实动作影响
  • 在5个真实场景任务中超越现有通用机器人模型表现
  • 无需训练,仅在测试时注入仿真,适合需要精细物理理解的任务

视觉语言模型(VLMs)具有出色的常识与语义推理能力,但缺乏对物理动态的具身理解。其训练数据为静态互联网级图文对,不含因果交互或动作相关的动态变化。因此,难以用于需要物理理解与精确动作规划的复杂机器人操作任务。为此,我们提出SIMPACT——一种测试时、基于仿真的动作规划框架,通过仿真闭环世界建模赋予VLM物理推理能力,无需额外训练。从单张RGB-D观测出发,SIMPACT高效构建物理仿真,使VLM可提出动作、观察模拟推演结果,并迭代优化决策。结合语言推理与物理预测,该方法能以具身方式理解接触动力学与动作后果。在五个需精细物理推理的真实刚体与柔性物体操作任务中,SIMPACT达到当前最优表现,优于现有通用机器人操作模型。结果表明,通过高效仿真嵌入物理理解至VLM推理中,是实现可泛化具身智能的可行路径。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation arises from training VLMs on static internet-scale visual-language data that contain no causal interactions or action-conditioned changes. Consequently, it remains challenging to leverage VLMs for fine-grained robotic manipulation tasks that require physical understanding, reasoning, and corresponding action planning. To overcome this, we present SIMPACT, a test-time, SIMulation-enabled ACTion Planning framework that equips VLMs with physical reasoning through simulation-in-the-loop world modeling, without requiring any additional training. From a single RGB-D observation, SIMPACT efficiently constructs physics simulations, enabling the VLM to propose informed actions, observe simulated rollouts, and iteratively refine its reasoning. By integrating language reasoning with physics prediction, our simulation-enabled VLM can understand contact dynamics and action outcomes in a physically grounded way. Our method demonstrates state-of-the-art performance on five challenging, real-world rigid-body and deformable manipulation tasks that require fine-grained physical reasoning, outperforming existing general-purpose robotic manipulation models. Our results demonstrate that embedding physics understanding via efficient simulation into VLM reasoning at test time offers a promising path towards generalizable embodied intelligence. Project webpage can be found at https://simpact-bot.github.io

机器人操作视觉语言模型物理仿真具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。