用不确定性引导的扩散模型生成分层强化学习子目标,提升采样效率与性能。
Hierarchical Reinforcement Learning with Uncertainty-Guided Diffusional Subgoals
- 用高斯过程正则化的条件扩散模型生成多样化子目标。
- 在连续控制基准上实现更高样本效率和更优性能。
- 适合研究分层强化学习与不确定性建模的学者参考。
分层强化学习(HRL)通过多时间抽象层次做出决策。其关键挑战在于低层策略随时间变化,使高层策略难以生成有效子目标。为此,高层策略需捕捉复杂的子目标分布,并考虑估计中的不确定性。本文提出一种方法:训练一个由高斯过程(GP)先验正则化的条件扩散模型,以生成多样化的子目标,并利用基于原理的GP不确定性量化。在此框架基础上,我们设计了一种策略,从扩散策略和GP预测均值中选择子目标。该方法在具有挑战性的连续控制基准测试中,显著优于现有HRL方法,在样本效率和性能方面均表现更优。
原文摘要 · Abstract (English)
Hierarchical reinforcement learning (HRL) learns to make decisions on multiple levels of temporal abstraction. A key challenge in HRL is that the low-level policy changes over time, making it difficult for the high-level policy to generate effective subgoals. To address this issue, the high-level policy must capture a complex subgoal distribution while also accounting for uncertainty in its estimates. We propose an approach that trains a conditional diffusion model regularized by a Gaussian Process (GP) prior to generate a complex variety of subgoals while leveraging principled GP uncertainty quantification. Building on this framework, we develop a strategy that selects subgoals from both the diffusion policy and GP's predictive mean. Our approach outperforms prior HRL methods in both sample efficiency and performance on challenging continuous control benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。