arXiv:2501.09080cs.LGcs.AI2025-01中稿 · the 2nd Reinforcem…被引 11

提出面向平均奖励的软演员-评论家算法,解决长期任务建模难题。

Average-Reward Soft Actor-Critic

  • 基于熵正则化设计平均奖励下的演员-评论家框架
  • 在标准基准上优于现有平均奖励算法,提升性能表现
  • 适合需要稳定长期奖励优化的研究者参考

近年来,平均奖励形式化强化学习因无需依赖折扣因子而受到关注,特别适用于解决时序延展性问题。在折扣设定下,带熵正则化的算法已显著优于确定性方法。然而,针对熵正则化平均奖励目标的深度强化学习算法仍缺乏。尽管近期已有基于策略梯度的平均奖励方法,但相应的演员-评论家框架尚未充分探索。本文提出一种平均奖励软演员-评论家算法,填补该空白。通过在标准强化学习基准上与现有平均奖励算法对比验证,所提方法在平均奖励指标上取得更优表现。

原文摘要 · Abstract (English)

The average-reward formulation of reinforcement learning (RL) has drawn increased interest in recent years for its ability to solve temporally-extended problems without relying on discounting. Meanwhile, in the discounted setting, algorithms with entropy regularization have been developed, leading to improvements over deterministic methods. Despite the distinct benefits of these approaches, deep RL algorithms for the entropy-regularized average-reward objective have not been developed. While policy-gradient based approaches have recently been presented for the average-reward literature, the corresponding actor-critic framework remains less explored. In this paper, we introduce an average-reward soft actor-critic algorithm to address these gaps in the field. We validate our method by comparing with existing average-reward algorithms on standard RL benchmarks, achieving superior performance for the average-reward criterion.

强化学习平均奖励熵正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。