四足机器人用视觉语言模型实现室内零样本物品抓取
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
- 结合前置抓手与视觉语言模型,实现环境语义理解
- 无需真实数据训练,60%成功率完成跨环境物品获取任务
- 适合对人机交互和复杂场景移动感兴趣的开发者
基于学习的方法在四足机器人运动方面表现优异,但其在需要与环境及人类互动的室内辅助技能方面仍面临挑战:缺乏操作末端执行器、仅依赖仿真数据导致语义理解有限、以及室内环境通行性与可达性差。本文提出一种面向室内环境的四足移动操作系统。系统配备前装抓手用于物体操作,采用基于第一视角深度图的仿真训练低层控制器,实现攀爬、全身倾斜等敏捷动作;同时利用预训练视觉语言模型(VLMs),结合第三人称鱼眼相机与第一视角RGB相机,实现语义理解与指令生成。我们在两个未见过的真实环境中评估该系统,未进行任何真实数据采集或训练。系统可零样本泛化至这些环境,并成功完成任务,如在爬过一张皇后尺寸床后按用户指令抓取随机放置的玩具,成功率达60%。
原文摘要 · Abstract (English)
Learning-based methods have achieved strong performance for quadrupedal locomotion. However, several challenges prevent quadrupeds from learning helpful indoor skills that require interaction with environments and humans: lack of end-effectors for manipulation, limited semantic understanding using only simulation data, and low traversability and reachability in indoor environments. We present a system for quadrupedal mobile manipulation in indoor environments. It uses a front-mounted gripper for object manipulation, a low-level controller trained in simulation using egocentric depth for agile skills like climbing and whole-body tilting, and pre-trained vision-language models (VLMs) with a third-person fisheye and an egocentric RGB camera for semantic understanding and command generation. We evaluate our system in two unseen environments without any real-world data collection or training. Our system can zero-shot generalize to these environments and complete tasks, like following user's commands to fetch a randomly placed stuff toy after climbing over a queen-sized bed, with a 60% success rate. Project website: https://helpful-doggybot.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。