让机器人理解复杂指令并实时响应反馈,实现开放式任务执行。
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

- 分层结构结合视觉、语言与动作模型,先推理再执行
- 能处理如'不要那个'等情境化反馈,支持动态调整
- 在三类机器人上验证,可完成清洁、做三明治等复杂任务
通用机器人在开放世界中执行多样化任务,需具备推理任务步骤、理解复杂指令及执行过程中的反馈能力。本文提出一种分层视觉-语言-动作系统,首先对复杂指令和用户反馈进行语义推理,确定最合适的下一步动作,再通过底层动作执行。相较于仅能完成简单命令(如“拿起杯子”)的直接指令跟随方法,本系统能处理如“给我做个素食三明治”或“那个不要”等包含上下文与反馈的复杂指令。我们在单臂、双臂及移动双臂三种机器人平台上进行了评估,验证了其在清理餐桌、制作三明治和杂货购物等任务中的有效性。相关视频见 https://www.pi.website/research/hirobot。
原文摘要 · Abstract (English)
Generalist robots that can perform a range of different tasks in open-world settings must be able to not only reason about the steps needed to accomplish their goals, but also process complex instructions, prompts, and even feedback during task execution. Intricate instructions (e.g., "Could you make me a vegetarian sandwich?" or "I don't like that one") require not just the ability to physically perform the individual steps, but the ability to situate complex commands and feedback in the physical world. In this work, we describe a system that uses vision-language models in a hierarchical structure, first reasoning over complex prompts and user feedback to deduce the most appropriate next step to fulfill the task, and then performing that step with low-level actions. In contrast to direct instruction following methods that can fulfill simple commands ("pick up the cup"), our system can reason through complex prompts and incorporate situated feedback during task execution ("that's not trash"). We evaluate our system across three robotic platforms, including single-arm, dual-arm, and dual-arm mobile robots, demonstrating its ability to handle tasks such as cleaning messy tables, making sandwiches, and grocery shopping. Videos are available at https://www.pi.website/research/hirobot
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。