arXiv:2505.12611cs.LG2025-05被引 3

提出新奖励设计方法,让智能体在困难环境中探索更有效且不偏离最优策略。

Action-Dependent Optimality-Preserving Reward Shaping

  • 基于动作的奖励转换机制,使内在激励可依赖行为选择
  • 在蒙特祖玛复仇者游戏中实现高效探索,胜过传统方法
  • 适合长期、稀疏奖励的复杂任务,尤其适用于深度强化学习

近期强化学习研究采用奖励塑形,特别是内在动机(IM)类复杂奖励,以促进稀疏奖励环境中的智能体探索。然而,奖励劫持可能导致内在奖励被过度优化而牺牲外在奖励,从而产生次优策略。基于潜在函数的奖励塑形(PBRS)方法如广义奖励匹配(GRM)和策略不变显式塑形(PIES)虽缓解此问题,但对长周期、高探索需求的复杂环境效果不佳。本文提出动作依赖型最优性保持塑形(ADOPS),将内在奖励转化为保持最优策略的形式,显著提升蒙特祖玛复仇者这类稀疏环境中的探索效率。我们证明ADOPS可处理非潜在形式的奖励函数:与要求累积内在回报与动作无关的PBRS不同,ADOPS允许其依赖于智能体行为,同时仍保持最优策略集合不变。实验表明,动作依赖性使ADOPS能在其他方法难以奏效的复杂稀疏环境中持续学习并维持最优性。

原文摘要 · Abstract (English)

Recent RL research has utilized reward shaping--particularly complex shaping rewards such as intrinsic motivation (IM)--to encourage agent exploration in sparse-reward environments. While often effective, ``reward hacking'' can lead to the shaping reward being optimized at the expense of the extrinsic reward, resulting in a suboptimal policy. Potential-Based Reward Shaping (PBRS) techniques such as Generalized Reward Matching (GRM) and Policy-Invariant Explicit Shaping (PIES) have mitigated this. These methods allow for implementing IM without altering optimal policies. In this work we show that they are effectively unsuitable for complex, exploration-heavy environments with long-duration episodes. To remedy this, we introduce Action-Dependent Optimality Preserving Shaping (ADOPS), a method of converting intrinsic rewards to an optimality-preserving form that allows agents to utilize IM more effectively in the extremely sparse environment of Montezuma's Revenge. We also prove ADOPS accommodates reward shaping functions that cannot be written in a potential-based form: while PBRS-based methods require the cumulative discounted intrinsic return be independent of actions, ADOPS allows for intrinsic cumulative returns to be dependent on agents' actions while still preserving the optimal policy set. We show how action-dependence enables ADOPS's to preserve optimality while learning in complex, sparse-reward environments where other methods struggle.

强化学习奖励塑形探索效率蒙特祖玛复仇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。