提出高效分层强化学习算法,实现高低层策略并行学习。
Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

- 通过低层多步转移特性设计并行学习机制
- 理论证明样本复杂度多项式增长,优于非分层方法
- 适合稀疏奖励下的目标导向任务,如机器人控制
我们提出 HBPI-UCRL,一种基于模型的分层强化学习(HRL)算法,可并行学习高层和低层策略。该算法利用高层转移对应低层多步转移的特性,引入两个关于低层动态的充分条件,使并行分层学习成为可能。当这些条件满足时,我们证明了 HBPI-UCRL 的样本复杂度在问题参数上为多项式。在稀疏奖励、目标导向设置下,其样本复杂度上界严格低于非分层版本,为分层强化学习的实证成功提供了理论支持。
原文摘要 · Abstract (English)
We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact that a high-level transition corresponds to a multi-step transition at the low level. We introduce two conditions on the low-level dynamics that are sufficient to make parallel HRL learnable. When these conditions hold, we prove that HBPI-UCRL has a polynomial sample complexity in the problem parameters. In the sparse-reward, goal-directed setting, our sample complexity upper bound for HBPI-UCRL is strictly lower than that of its non-hierarchical counterpart, providing theoretical justification for the empirical success of HRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。