arXiv:2510.07778cs.ROcs.AI2025-10被引 4

让机器人理解人类隐含意图,提升人机交互成功率。

IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction

  • 用分阶段训练增强模型推理与感知能力,融合意图推断和空间定位。
  • 直接指令下成功率达18%提升,间接指令下超基线28%。
  • 零样本交互达40%成功率,适合复杂真实场景的智能机器人。

视觉-语言-动作(VLA)模型通过预训练视觉-语言模型将感知与机器人控制结合,为通用具身智能提供新路径。但现有最先进方法主要在与具身场景关联有限的多模态任务上预训练,再微调以实现显式指令到动作的映射,缺乏推理密集型预训练与推理引导的操作,难以处理复杂现实交互中所需的隐含意图推理。为此,我们提出IntentionVLA,采用课程训练范式与高效推理机制。首先利用精心设计的推理数据集,融合意图推断、空间定位与紧凑具身推理,赋予模型推理与感知双重能力;随后在微调阶段,以紧凑推理输出作为动作生成的上下文指导,实现间接指令下的快速推理。实验表明,IntentionVLA显著优于π₀,在直接指令下成功率达18%提升,间接指令下较ECoT高28%。在分布外意图任务中,成功率达所有基线两倍以上,并实现40%零样本人机交互成功率。这些结果凸显IntentionVLA作为下一代人机交互系统的潜力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose embodied intelligence. However, current SOTA VLAs are primarily pretrained on multimodal tasks with limited relevance to embodied scenarios, and then finetuned to map explicit instructions to actions. Consequently, due to the lack of reasoning-intensive pretraining and reasoning-guided manipulation, these models are unable to perform implicit human intention reasoning required for complex, real-world interactions. To overcome these limitations, we propose \textbf{IntentionVLA}, a VLA framework with a curriculum training paradigm and an efficient inference mechanism. Our proposed method first leverages carefully designed reasoning data that combine intention inference, spatial grounding, and compact embodied reasoning, endowing the model with both reasoning and perception capabilities. In the following finetuning stage, IntentionVLA employs the compact reasoning outputs as contextual guidance for action generation, enabling fast inference under indirect instructions. Experimental results show that IntentionVLA substantially outperforms $π_0$, achieving 18\% higher success rates with direct instructions and 28\% higher than ECoT under intention instructions. On out-of-distribution intention tasks, IntentionVLA achieves over twice the success rate of all baselines, and further enables zero-shot human-robot interaction with 40\% success rate. These results highlight IntentionVLA as a promising paradigm for next-generation human-robot interaction (HRI) systems.

人机交互意图推理具身智能VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。