通过调整势函数提升强化学习中的奖励塑造效果
Improving the Effectiveness of Potential-Based Reward Shaping in Reinforcement Learning
- 用线性平移势函数增强奖励塑造有效性
- 在不改变初始Q值前提下显著提升样本效率
- 适用于稀疏奖励场景,适合深度强化学习研究者
基于势函数的奖励塑造常用于将任务先验知识融入强化学习,因其能保证策略不变性。本文指出,有效奖励塑造依赖于初始Q值和外部奖励,这影响了智能体利用塑造奖励引导探索的能力。我们形式化推导出:通过简单线性平移势函数,可在不改变其编码偏好、无需调整初始Q值的前提下,提升塑造效果。我们揭示了连续势函数在正确分配正负塑造奖励上的理论局限。在网格世界、倒立摆和山地车环境中的实验验证了理论结果,表明该方法在稀疏奖励设置下可有效提升深度强化学习的样本效率。
原文摘要 · Abstract (English)
Potential-based reward shaping is commonly used to incorporate prior knowledge of how to solve the task into reinforcement learning because it can formally guarantee policy invariance. As such, the optimal policy and the ordering of policies by their returns are not altered by potential-based reward shaping. In this work, we highlight the dependence of effective potential-based reward shaping on the initial Q-values and external rewards, which determine the agent's ability to exploit the shaping rewards to guide its exploration and achieve increased sample efficiency. We formally derive how a simple linear shift of the potential function can be used to improve the effectiveness of reward shaping without changing the encoded preferences in the potential function, and without having to adjust the initial Q-values, which can be challenging and undesirable in deep reinforcement learning. We show the theoretical limitations of continuous potential functions for correctly assigning positive and negative reward shaping values. We verify our theoretical findings empirically on Gridworld domains with sparse and uninformative reward functions, as well as on the Cart Pole and Mountain Car environments, where we demonstrate the application of our results in deep reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。