用人类实时干预提升机器人学习效率,15分钟完成训练。
Data-Efficient Learning from Human Interventions for Mobile Robots
- 结合模仿与强化学习,通过在线人类干预实时优化策略
- 无需奖励函数或预训练,仅用15分钟完成两台机器人的任务训练
- 适合需要快速安全部署的现实机器人场景
移动机器人在自动驾驶配送和服务业中至关重要。基于学习的方法因其鲁棒性和泛化能力受到青睐,但传统模仿学习(IL)和强化学习(RL)需大量数据、精心设计的奖励函数,并存在仿真到现实的差距,难以高效安全地部署。我们提出一种在线人机协同学习方法 PVP4Real,融合 IL 与 RL,实现从实时人类干预与示范中高效学习,无需奖励函数或预训练,显著提升数据效率与训练安全性。我们在两种移动机器人任务中验证该方法:一只四足机器人和一辆轮式配送机器人,其中一项任务使用原始 RGBD 图像作为观测。所有训练均在15分钟内完成。实验表明,人机协同学习有望解决真实机器人任务中的数据效率问题。
原文摘要 · Abstract (English)
Mobile robots are essential in applications such as autonomous delivery and hospitality services. Applying learning-based methods to address mobile robot tasks has gained popularity due to its robustness and generalizability. Traditional methods such as Imitation Learning (IL) and Reinforcement Learning (RL) offer adaptability but require large datasets, carefully crafted reward functions, and face sim-to-real gaps, making them challenging for efficient and safe real-world deployment. We propose an online human-in-the-loop learning method PVP4Real that combines IL and RL to address these issues. PVP4Real enables efficient real-time policy learning from online human intervention and demonstration, without reward or any pretraining, significantly improving data efficiency and training safety. We validate our method by training two different robots -- a legged quadruped, and a wheeled delivery robot -- in two mobile robot tasks, one of which even uses raw RGBD image as observation. The training finishes within 15 minutes. Our experiments show the promising future of human-in-the-loop learning in addressing the data efficiency issue in real-world robotic tasks. More information is available at: https://metadriverse.github.io/pvp4real/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。