通过回溯目标生成动作正则,提升强化学习样本效率。
GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning
- 基于回溯目标生成动作正则化先验,增强经验利用。
- 在导航与操作任务中显著提升样本效率和性能。
- 适合追求高效强化学习的科研与工程人员。
目标条件强化学习(GCRL)在稀疏奖励场景下仍是核心挑战。尽管事后经验回放(HER)通过重标注轨迹中的达成目标展现潜力,但仅靠轨迹重标注无法充分挖掘离策略GCRL方法中的可用经验,导致样本效率受限。本文提出事后目标条件正则化(HGR),基于回溯目标生成动作正则化先验。结合事后自模仿正则化(HSR),该方法使离策略强化学习算法最大化经验利用率。在一系列导航与操作任务上,相较于使用HER与自模仿技术的现有方法,本方法实现更高效的样本复用与最优性能。
原文摘要 · Abstract (English)
Goal-conditioned reinforcement learning (GCRL) with sparse rewards remains a fundamental challenge in reinforcement learning. While hindsight experience replay (HER) has shown promise by relabeling collected trajectories with achieved goals, we argue that trajectory relabeling alone does not fully exploit the available experiences in off-policy GCRL methods, resulting in limited sample efficiency. In this paper, we propose Hindsight Goal-conditioned Regularization (HGR), a technique that generates action regularization priors based on hindsight goals. When combined with hindsight self-imitation regularization (HSR), our approach enables off-policy RL algorithms to maximize experience utilization. Compared to existing GCRL methods that employ HER and self-imitation techniques, our hindsight regularizations achieve substantially more efficient sample reuse and the best performances, which we empirically demonstrate on a suite of navigation and manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。