arXiv:2608.08839cs.ROcs.CV2026-08

用视觉语言模型提升机器人动作模型的指令理解力

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

论文配图:SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models
图 1 · 摘自论文原文
  • 引入视觉语言模型生成带语义引导的未来场景预测
  • 同时捕捉目标物体与空间结构,确保动作精准
  • 在仿真和真实场景中均实现高精度指令跟随

世界-动作模型(WAMs)在机器人操作任务中展现出巨大潜力,但现有方法主要依赖视觉线索生成未来视频和动作,缺乏对语言指令的有效融合,因通用文本编码器独立于视觉输入。这导致预测结果与指令语义不一致,影响动作准确性。为此,我们提出SG-WAM,通过视觉语言模型(VLM)作为语义规划器,增强WAM的指令对齐能力。具体地,训练一个基于VLM的规划器,预测兼具文本锚定与空间感知的语义前景:前者定位正确目标物体,后者提供场景几何信息以支持精确操作。将此语义前景注入世界-动作模型,作为高层语义引导,使未来视频生成与动作预测严格遵循语言指令。大量仿真与真实世界实验表明,该方法显著提升了操作精度与指令遵循能力。

原文摘要 · Abstract (English)

World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that leverages a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of world-action models. Specifically, we train a VLM-based planner to predict text-grounded and spatial-aware semantic foresight. The text-grounded semantic foresight grounds the instruction by identifying the correct target objects, and the spatial-aware semantic foresight provides the scene geometry for precise manipulation. We then inject this foresight into the world-action model as high-level semantic guidance, ensuring that both future-video generation and action prediction faithfully follow the language instruction. Extensive experiments in simulation and the real world demonstrate the superiority of our semantic guidance method, showcasing precise manipulation and strong instruction-following capabilities.

机器人操作视觉语言模型指令跟随语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。