用AR远程交互收集人类示范数据,提升机械臂灵巧操作学习效率
Scalable Dexterous Robot Learning with AR-based Remote Human-Robot Interactions
- 通过AR远程交互采集专家示范数据,用于行为克隆预训练
- 引入对比学习增强的强化学习,成功率显著高于基线方法
- 适用于需要高效、安全灵巧操作的工业机器人场景
本文研究灵巧机械臂系统中可扩展的机器人操作学习问题,通过基于增强现实(AR)的远程人机交互系统收集专家示范数据以提高学习效率。提出两阶段方法:第一阶段采用行为克隆(BC)方式利用AR系统生成的学习数据进行策略预训练;第二阶段设计一种基于对比学习的强化学习(RL)方法,引入投影头加速学习进程,并采用事件驱动的增强奖励机制提升安全性。在PyBullet物理仿真和真实世界实验中验证,相比基线方法,该方法不仅显著加快训练速度,且在任务成功率上表现更优。消融实验证明,对比学习有效避免了策略坍塌现象。补充演示视频见https://cyberyyc.github.io/。
原文摘要 · Abstract (English)
This paper focuses on the scalable robot learning for manipulation in the dexterous robot arm-hand systems, where the remote human-robot interactions via augmented reality (AR) are established to collect the expert demonstration data for improving efficiency. In such a system, we present a novel method to address the general manipulation task problem. Specifically, the proposed method consists of two phases: i) In the first phase for pretraining, the policy is created in a behavior cloning (BC) manner, through leveraging the learning data from our AR-based remote human-robot interaction system; ii) In the second phase, a contrastive learning empowered reinforcement learning (RL) method is developed to obtain more efficient and robust policy than the BC, and thus a projection head is designed to accelerate the learning progress. An event-driven augmented reward is adopted for enhancing the safety. To validate the proposed method, both the physics simulations via PyBullet and real-world experiments are carried out. The results demonstrate that compared to the baselines, our method not only significantly speeds up the training process, but also achieves much better performance in terms of the success rate for fulfilling the manipulation tasks. By conducting the ablation study, it is confirmed that the proposed RL with contrastive learning overcomes policy collapse. Supplementary demonstrations are available at https://cyberyyc.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。