机器人通过检索人类视频示范学习新任务,提升泛化能力。
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- 从人类视频中提取物体功能掩码和手部轨迹,构建知识库。
- 根据任务描述检索相关视频,显著提升复杂场景下的成功率。
- 适合需要快速适应新任务的现实机器人系统。
在复杂不确定环境中,机器人面临巨大挑战。现有先进系统依赖大规模数据集进行训练,而人类面对陌生任务时,常通过观看视频示范来学习。本文提出一种基于视频检索(RfV)的机器人策略学习方法,借鉴人类示范机制解决操作任务。系统构建包含多样日常任务的人类视频数据库,并提取物体功能掩码、手部运动轨迹等中层信息作为增强输入,以提升模型学习与泛化能力。设计双组件架构:视频检索器从外部视频库中根据任务描述获取相关视频,策略生成器将检索知识融入学习循环。该方法使机器人能够针对不同场景生成自适应响应,并推广至训练数据外的任务。在多个模拟与真实场景中测试表明,本系统性能显著优于传统方法,是机器人学习领域的重要突破。
原文摘要 · Abstract (English)
Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as assembling a chair, a common approach is to learn by watching video demonstrations. In this paper, we propose a novel method for learning robot policies by Retrieving-from-Video (RfV), using analogies from human demonstrations to address manipulation tasks. Our system constructs a video bank comprising recordings of humans performing diverse daily tasks. To enrich the knowledge from these videos, we extract mid-level information, such as object affordance masks and hand motion trajectories, which serve as additional inputs to enhance the robot model's learning and generalization capabilities. We further feature a dual-component system: a video retriever that taps into an external video bank to fetch task-relevant video based on task specification, and a policy generator that integrates this retrieved knowledge into the learning cycle. This approach enables robots to craft adaptive responses to various scenarios and generalize to tasks beyond those in the training data. Through rigorous testing in multiple simulated and real-world settings, our system demonstrates a marked improvement in performance over conventional robotic systems, showcasing a significant breakthrough in the field of robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。