提出新方法让智能体探索时不偏离最优策略,避免盲目试错。
Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards
- 用潜在函数重构内在奖励,确保最优策略不变
- 在门钥匙和悬崖行走任务中避免陷入次优解,加速训练
- 适合需安全探索的强化学习场景,如机器人控制
近期涌现大量内在动机(IM)奖励塑造方法,用于复杂稀疏奖励环境中的学习。然而这些方法常无意改变最优策略集合,导致次优行为。传统基于势能的奖励塑造(PBRS)难以适用于多数复杂的可训练IM方法,因其依赖更广泛变量。本文提出PBRS的扩展,证明其在更广函数类下仍保持最优策略集合不变,并引入潜在基内在动机(PBIM)与广义奖励匹配(GRM),将IM奖励转化为势能形式,不改变最优策略。在MiniGrid DoorKey和Cliff Walking环境中验证,PBIM与GRM有效防止智能体收敛至次优策略,并加快训练速度。此外证明GRM足够通用,可涵盖所有基于势能的奖励塑造函数。本文扩展了PBIM方法,提出更通用的GRM,补充额外证明、实验结果与讨论。
原文摘要 · Abstract (English)
Recently there has been a proliferation of intrinsic motivation (IM) reward-shaping methods to learn in complex and sparse-reward environments. These methods can often inadvertently change the set of optimal policies in an environment, leading to suboptimal behavior. Previous work on mitigating the risks of reward shaping, particularly through potential-based reward shaping (PBRS), has not been applicable to many IM methods, as they are often complex, trainable functions themselves, and therefore dependent on a wider set of variables than the traditional reward functions that PBRS was developed for. We present an extension to PBRS that we prove preserves the set of optimal policies under a more general set of functions than has been previously proven. We also present {\em Potential-Based Intrinsic Motivation} (PBIM) and {\em Generalized Reward Matching} (GRM), methods for converting IM rewards into a potential-based form that are useable without altering the set of optimal policies. Testing in the MiniGrid DoorKey and Cliff Walking environments, we demonstrate that PBIM and GRM successfully prevent the agent from converging to a suboptimal policy and can speed up training. Additionally, we prove that GRM is sufficiently general as to encompass all potential-based reward shaping functions. This paper expands on previous work introducing the PBIM method, and provides an extension to the more general method of GRM, as well as additional proofs, experimental results, and discussion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。