提出一种新Q-learning算法,解决长期奖励下的风险规避问题。
Risk-Averse Total-Reward Reinforcement Learning
- 基于动态一致性与可提取性,设计风险规避型Q-learning
- 在小规模表格环境中快速收敛至最优风险规避值函数
- 适合需稳健长期决策的强化学习应用
风险规避的总奖励马尔可夫决策过程(MDP)为建模和求解无折扣无限时域目标提供了有前景的框架。现有的基于模型的算法在处理如熵风险度量(ERM)和熵值风险(EVaR)等风险度量时表现有效,但需要完全访问转移概率。本文提出一种Q-learning算法,用于计算总奖励下ERM和EVaR目标的最优平稳策略,并具备强收敛性和性能保证。算法及其最优性得益于ERM的动态一致性和可提取性。在表格化领域中的数值结果表明,所提出的Q-learning算法能快速且可靠地收敛到最优风险规避值函数。
原文摘要 · Abstract (English)
Risk-averse total-reward Markov Decision Processes (MDPs) offer a promising framework for modeling and solving undiscounted infinite-horizon objectives. Existing model-based algorithms for risk measures like the entropic risk measure (ERM) and entropic value-at-risk (EVaR) are effective in small problems, but require full access to transition probabilities. We propose a Q-learning algorithm to compute the optimal stationary policy for total-reward ERM and EVaR objectives with strong convergence and performance guarantees. The algorithm and its optimality are made possible by ERM's dynamic consistency and elicitability. Our numerical results on tabular domains demonstrate quick and reliable convergence of the proposed Q-learning algorithm to the optimal risk-averse value function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。