提出新方法解决非稳态环境下持续任务的平均奖励学习问题
A Harmonic Mean Formulation of Average Reward Reinforcement Learning in SMDPs

- 用改进的调和平均算子替代传统奖赏/时长比计算
- 在非稳态分布下仍能准确估计长期平均奖励率
- 适用于持续运行的复杂决策任务,适合强化学习研究者
近期研究重新激发了对无限时域、非周期性(持续)任务中无折扣平均奖励强化学习算法的兴趣。半马尔可夫决策过程(SMDPs)尤为关键:离散动作随机产生奖励与持续时间,目标是优化平均奖励速率。现有方法通过优化奖励与持续时间之比实现,但在奖励与持续时间非平稳(无限时域)时可能失效。本文提出一种新型修正调和平均算子,可在该条件下正确计算奖励速率。由此导出无需模型的算法,能在SMDPs中有效运行,且对随时间变化的奖励与持续时间分布保持鲁棒性。我们证明了该算子的理论性质,并通过实验证明其优于现有算法。
原文摘要 · Abstract (English)
Recent research has revived and amplified interest in algorithms for undiscounted average reward reinforcement learning in infinite-horizon, non-episodic (continuing) tasks. Semi-Markov decision processes (SMDPs) are of particular interest. In SMDPs, discrete actions stochastically generate both rewards and durations, and the objective is to optimize the average reward rate. Existing algorithms approach this by optimizing the ratio of rewards to durations. However, when rewards and durations are non-stationary (in the infinite horizon), this can be incorrect. This paper presents a novel modified harmonic mean operator that correctly computes reward rates even under such conditions. This yields model-free learning algorithms that can work with SMDPs, while maintaining robustness to non-stationary reward and duration distributions over time. We prove theoretical properties of the modified harmonic mean operator, and empirically demonstrate its efficacy in comparison to existing algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。