arXiv:2505.13888cs.RO2025-05被引 21

让机器人更懂空间关系,提升指令理解的准确性。

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

  • 在指令前加空间方向问题,引导模型关注关键视觉信息。
  • 实测在仿真与真实环境均显著提升动作执行准确率。
  • 无需额外训练,可直接插件式增强现有机器人模型。

利用预训练视觉语言模型(VLMs)将语言指令与视觉观测映射为原始低层动作,视觉-语言-动作模型(VLAs)有望实现通用机器人系统。然而,现有VLAs易将任务无关的视觉特征与动作错误关联,限制其泛化能力。为此,我们提出内在空间推理(InSpire),通过增强VLAs的空间推理能力来缓解虚假相关性。具体地,InSpire在语言指令前添加问题:“[物体]相对于机器人在哪个方向?”,并使答案(左/右/上/下/前/后/已抓取)与预测动作对齐真实标注。值得注意的是,InSpire可作为插件直接提升现有自回归VLAs,无需额外训练数据或与其他大模型交互。大量仿真与真实环境实验表明该方法有效且灵活。

原文摘要 · Abstract (English)

Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correlate task-irrelevant visual features with actions, limiting their generalization capacity beyond the training data. To tackle this challenge, we propose Intrinsic Spatial Reasoning (InSpire), a simple yet effective approach that mitigates the adverse effects of spurious correlations by boosting the spatial reasoning ability of VLAs. Specifically, InSpire redirects the VLA's attention to task-relevant factors by prepending the question "In which direction is the [object] relative to the robot?" to the language instruction and aligning the answer "right/left/up/down/front/back/grasped" and predicted actions with ground-truth. Notably, InSpire can be used as a plugin to enhance existing autoregressive VLAs, requiring no extra training data or interaction with other large models. Extensive experimental results in both simulation and real-world environments demonstrate the effectiveness and flexibility of our approach.

机器人空间推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。