用真实视频训练机器人按语言指令找物,速度快且效果更好
LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos
- 用大模型自动标注无标签真实场景视频,实现零人工干预
- 130小时多源数据训练,1000+实测验证性能超越现有方法
- 推理速度达竞品4倍,适合部署在边缘设备的机器人导航
世界充满各类物体。为让机器人实用,需具备根据人类描述找到任意物体的能力。本文提出LeLaN(学习语言条件导航策略),一种利用无标签、无动作的视角内数据学习可扩展语言条件物体导航的新方法。该框架借助大规模视觉-语言模型和机器人基础模型,对来自室内和室外多种环境的真实数据进行自动标注,涵盖机器人观测、YouTube旅游视频和人类步行数据,共标注超过130小时数据。在超过1000次真实测试中,该方法证明仅凭无标签视频即可训练出性能超越当前最优机器人导航方案的策略,同时在边缘计算设备上推理速度达到其4倍。项目代码、数据集及补充视频已开源。
原文摘要 · Abstract (English)
The world is filled with a wide variety of objects. For robots to be useful, they need the ability to find arbitrary objects described by people. In this paper, we present LeLaN(Learning Language-conditioned Navigation policy), a novel approach that consumes unlabeled, action-free egocentric data to learn scalable, language-conditioned object navigation. Our framework, LeLaN leverages the semantic knowledge of large vision-language models, as well as robotic foundation models, to label in-the-wild data from a variety of indoor and outdoor environments. We label over 130 hours of data collected in real-world indoor and outdoor environments, including robot observations, YouTube video tours, and human walking data. Extensive experiments with over 1000 real-world trials show that our approach enables training a policy from unlabeled action-free videos that outperforms state-of-the-art robot navigation methods, while being capable of inference at 4 times their speed on edge compute. We open-source our models, datasets and provide supplementary videos on our project page (https://learning-language-navigation.github.io/).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。