通过分步推理物体交互能力,提升机器人在复杂任务中的表现
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
- 用四个层次的交互认知引导动作决策:物体、抓取、空间、移动
- 在多个任务上超越OpenVLA和Octo,对未见物体姿态和新环境有强泛化能力
- 将视觉与文本提示融合,让模型在动作推断中利用上下文信息
机器人基础模型,尤其是视觉-语言-动作(VLA)模型,因其显著提升机器人策略学习的泛化与鲁棒性而备受关注。OpenAI最新模型O1通过长链推理展现强大解决复杂问题的能力。这引发一个关键问题:机器人能否通过回顾先前观察,并生成任务特定的推理,来指导动作预测?本文提出链式可及性(CoA-VLA),一种通过引入序列化机器人可及性推理来扩展机器人模型的新方法。具体而言,模型在行动前需考虑四类可及性:(1) 物体可及性——操作哪个物体及其位置;(2) 抓取可及性——抓取物体的特定部位;(3) 空间可及性——放置物体的最佳空间;(4) 移动可及性——无碰撞的运动路径。我们将每类可及性转化为视觉与文本两种提示形式,并设计新型视觉-语言共注入模块,将这些知识融入策略网络。该机制使机器人在动作推理中有效利用上下文信息,提升精度与鲁棒性。实验表明,CoA-VLA在多种任务上优于当前最优机器人基础模型(如OpenVLA与Octo),并展现出强泛化能力,包括识别未见物体姿态、定位自由空间及在新环境中避障。
原文摘要 · Abstract (English)
Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent model, O1, showcased impressive capabilities in solving complex problems by utilizing extensive reasoning chains. This prompts an important question: can robot models achieve better performance in multi-task , complex environments by reviewing prior observations and then providing task-specific reasoning to guide action prediction? In this paper, we introduce Chain-of-Affordance (CoA-VLA) , a novel approach to scaling robot models by incorporating reasoning in the format of sequential robot affordances to facilitate task completion. Specifically, we prompt the model to consider the following four types of affordances before taking action: (1) object affordance - what object to manipulate and where it is ; (2) grasp affordance - the specific object part to grasp ; (3) spatial affordance - the optimal space to place the object ; and (4) movement affordance-the collision - free path for movement. We further transform each affordance into two prompting formats: visual affordance and textual affordance. We introduce a novel vision-language co-injection module that integrates this knowledge into the policy network. This allows the robot to leverage essential contextual information during action inference, resulting in improved precision and robustness. Our experiments demonstrate that CoA-VLA outperforms state-of-the-art robot foundation models, including OpenVLA and Octo, on a variety of tasks. Furthermore, CoA-VLA exhibits strong generalization capabilities, including recognizing unseen object poses, identifying free space, and avoiding obstacles in novel environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。