首次将分布强化学习拓展到平均奖励场景,可学习长期每步收益分布。
A Differential Perspective on Distributional Reinforcement Learning
- 基于分位数方法构建平均奖励下的分布学习算法
- 提出收敛性证明的表格算法,支持预测与控制任务
- 性能优于传统方法,且能捕捉丰富收益分布信息
迄今为止,分布强化学习(distributional RL)方法均聚焦于折扣回报设置,即优化随时间累积的折扣奖励总和。本文将分布强化学习拓展至平均奖励设置,目标是优化每步获得的长期平均奖励。我们采用分位数方法,开发出首个能够成功学习和/或优化长期每步奖励分布及平均奖励马尔可夫决策过程(MDP)的差分回报分布的算法。我们推导了适用于预测与控制任务的收敛性证明的表格算法,并提出一组具有良好扩展特性的更广泛算法族。实验表明,这些算法在性能上与非分布式对应方法相当甚至更优,同时能有效捕捉长期每步奖励和差分回报分布的丰富信息。
原文摘要 · Abstract (English)
To date, distributional reinforcement learning (distributional RL) methods have exclusively focused on the discounted setting, where an agent aims to optimize a discounted sum of rewards over time. In this work, we extend distributional RL to the average-reward setting, where an agent aims to optimize the reward received per time step. In particular, we utilize a quantile-based approach to develop the first set of algorithms that can successfully learn and/or optimize the long-run per-step reward distribution, as well as the differential return distribution of an average-reward MDP. We derive proven-convergent tabular algorithms for both prediction and control, as well as a broader family of algorithms that have appealing scaling properties. Empirically, we find that these algorithms yield competitive and sometimes superior performance when compared to their non-distributional equivalents, while also capturing rich information about the long-run per-step reward and differential return distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。