用专家示范和任意重置模拟器加速在线强化学习
Accelerated Online Reinforcement Learning using Auxiliary Start State Distributions
- 引入辅助起始状态分布,不依赖真实起始分布
- 结合安全性和轨迹长度信息,显著提升样本效率
- 适合有专家数据或模拟器的复杂环境探索任务
在线强化学习中的长期难题是样本效率低下,根源在于无法高效探索环境。现有方法多从零开始学习,未能利用专家示范和可任意重置的模拟器。本文探索如何借助少量专家示范和具备任意状态重置能力的模拟器加速在线强化学习。研究发现,采用与马尔可夫决策过程真实起始分布不同的辅助起始状态分布,能显著提升样本效率。通过以安全性为指导选择该分布,并利用轨迹长度信息进行操作化,我们在稀疏奖励的高难度探索环境中实现了当前最优的样本效率。
原文摘要 · Abstract (English)
A long-standing problem in online reinforcement learning (RL) is of ensuring sample efficiency, which stems from an inability to explore environments efficiently. Most attempts at efficient exploration tackle this problem in a setting where learning begins from scratch, without prior information available to bootstrap learning. However, such approaches fail to leverage expert demonstrations and simulators that can reset to arbitrary states. These affordances are valuable resources that offer enormous potential to guide exploration and speed up learning. In this paper, we explore how a small number of expert demonstrations and a simulator allowing arbitrary resets can accelerate learning during online RL. We find that training with a suitable choice of an auxiliary start state distribution that may differ from the true start state distribution of the underlying Markov Decision Process can significantly improve sample efficiency. We find that using a notion of safety to inform the choice of this auxiliary distribution significantly accelerates learning. By using episode length information as a way to operationalize this notion, we demonstrate state-of-the-art sample efficiency on a sparse-reward hard-exploration environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。