arXiv:2604.11751cs.ROcs.AI2026-04被引 1

用语言指令指导机器人规划,实现跨场景通用决策。

Grounded World Model for Semantically Generalizable Planning

论文配图:Grounded World Model for Semantically Generalizable Planning
图 1 · 摘自论文原文
  • 在视觉-语言对齐空间中构建世界模型,用语义相似度评分动作效果。
  • 在新环境任务中达87%成功率,显著超越传统方法的22%。
  • 适合需要理解自然语言指令的复杂交互场景研究者。

在模型预测控制(MPC)中,世界模型预测不同动作提案的未来结果,并通过得分选择最优动作。对于视觉-运动型MPC,评分函数是预测图像与目标图像在预训练视觉编码器(如DINO和JEPA)的隐空间中的距离。然而,在任务执行前难以获取目标图像,尤其在新环境中;且仅通过图像传达目标互动性有限。本文提出在视觉-语言对齐的隐空间中学习一个接地世界模型(GWM)。每个动作的未来结果根据其与任务指令嵌入的相似度进行评分,从而将视觉-运动MPC转化为视觉-语言动作模型(VLA),在语义泛化上优于现有VLM-based VLA。在提出的WISER基准测试中,GWM-MPC在包含288个任务的测试集上取得87%的成功率,这些任务包含未见过的视觉信号和指代表达,但可通过训练中演示的动作解决。相比之下,传统VLA即使在训练集上达到90%成功率,仍仅维持22%的平均测试成功率。

原文摘要 · Abstract (English)

In Model Predictive Control (MPC), world models predict the future outcomes of various action proposals, which are then scored to guide the selection of the optimal action. For visuomotor MPC, the score function is a distance metric between a predicted image and a goal image, measured in the latent space of a pretrained vision encoder like DINO and JEPA. However, it is challenging to obtain the goal image in advance of the task execution, particularly in new environments. Additionally, conveying the goal through an image offers limited interactivity compared with natural language. In this work, we propose to learn a Grounded World Model (GWM) in a vision-language-aligned latent space. As a result, each proposed action is scored based on how close its future outcome is to the task instruction, reflected by the similarity of embeddings. This approach transforms the visuomotor MPC to a VLA that surpasses VLM-based VLAs in semantic generalization. On the proposed WISER benchmark, GWM-MPC achieves a 87% success rate on the test set comprising 288 tasks that feature unseen visual signals and referring expressions, yet remain solvable with motions demonstrated during training. In contrast, traditional VLAs achieve an average success rate of 22%, even though they overfit the training set with a 90% success rate.

视觉-语言机器人规划语义泛化世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。