用在线示范提升机器人策略迁移效果,更省样本且更稳定。
Robot Policy Transfer with Online Demonstrations: An Active Reinforcement Learning Approach
- 主动选择最优时机和内容获取在线示范,优化示范效率。
- 在8个场景中平均成功率显著高于基线方法,样本效率提升明显。
- 适合需要快速适应新任务的机器人系统,尤其擅长跨环境迁移。
迁移学习(TL)是使机器人在不同环境、任务或机械结构间转移已学策略的强大工具。为进一步促进该过程,已有研究尝试结合示范学习(LfD)以实现更灵活高效的策略迁移。然而,现有方法几乎仅依赖转移前收集的离线示范,易受示范偏差(covariance shift)影响,损害迁移性能。相比之下,在从零训练的设置中,在线示范已被证明可有效缓解此问题并提升样本效率。本文将这一洞察引入策略迁移场景,提出基于在线示范的主动示范算法——政策迁移与在线示范(Policy Transfer with Online Demonstrations)。该方法在有限示范预算下,动态优化在线示范的时机与内容。我们在8个机器人场景中评估该方法,涵盖不同环境特征、任务目标及机械结构间的策略迁移,旨在将源任务训练好的策略迁移到相关但不同的目标任务。结果表明,该方法在平均成功率和样本效率方面均显著优于两种使用离线示范的典型方法以及一种使用在线示范的主动示范方法。此外,我们在三个迁移场景中进行了初步的仿真到现实测试,验证了所迁移策略在真实机器人操作臂上的有效性。
原文摘要 · Abstract (English)
Transfer Learning (TL) is a powerful tool that enables robots to transfer learned policies across different environments, tasks, or embodiments. To further facilitate this process, efforts have been made to combine it with Learning from Demonstrations (LfD) for more flexible and efficient policy transfer. However, these approaches are almost exclusively limited to offline demonstrations collected before policy transfer starts, which may suffer from the intrinsic issue of covariance shift brought by LfD and harm the performance of policy transfer. Meanwhile, extensive work in the learning-from-scratch setting has shown that online demonstrations can effectively alleviate covariance shift and lead to better policy performance with improved sample efficiency. This work combines these insights to introduce online demonstrations into a policy transfer setting. We present Policy Transfer with Online Demonstrations, an active LfD algorithm for policy transfer that can optimize the timing and content of queries for online episodic expert demonstrations under a limited demonstration budget. We evaluate our method in eight robotic scenarios, involving policy transfer across diverse environment characteristics, task objectives, and robotic embodiments, with the aim to transfer a trained policy from a source task to a related but different target task. The results show that our method significantly outperforms all baselines in terms of average success rate and sample efficiency, compared to two canonical LfD methods with offline demonstrations and one active LfD method with online demonstrations. Additionally, we conduct preliminary sim-to-real tests of the transferred policy on three transfer scenarios in the real-world environment, demonstrating the policy effectiveness on a real robot manipulator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。