提出风险敏感的平均代价MDP规划与学习方法,让智能体更精细地应对不确定性。
Planning and Learning in Average Risk-aware MDPs
- 用相对值迭代和多层蒙特卡洛方法扩展传统算法以支持动态风险度量
- 理论证明算法收敛,实验验证可学习出对风险敏感的优化策略
- 适合需要权衡风险与收益的连续决策场景,如金融或自动驾驶
在持续任务中,平均代价马尔可夫决策过程(Average Cost MDP)具有明确价值,并可通过高效算法求解。然而,其默认假设为智能体完全风险中性。本文将风险中性算法推广至更一般的动态风险度量框架。具体提出一种用于规划的相对值迭代(RVI)算法,以及两种无模型的Q-learning算法:基于多层蒙特卡洛(MLMC)的通用算法,和针对基于效用的缺口风险度量设计的离策略算法。证明了RVI与MLMC-Q-learning均能收敛至最优。数值实验验证了分析结果,实证确认了离策略算法的收敛性,并表明本方法能有效识别出与智能体复杂风险偏好高度匹配的策略。
原文摘要 · Abstract (English)
For continuing tasks, average cost Markov decision processes have well-documented value and can be solved using efficient algorithms. However, it explicitly assumes that the agent is risk-neutral. In this work, we extend risk-neutral algorithms to accommodate the more general class of dynamic risk measures. Specifically, we propose a relative value iteration (RVI) algorithm for planning and design two model-free Q-learning algorithms, namely a generic algorithm based on the multi-level Monte Carlo (MLMC) method, and an off-policy algorithm dedicated to utility-based shortfall risk measures. Both the RVI and MLMC-based Q-learning algorithms are proven to converge to optimality. Numerical experiments validate our analysis, confirm empirically the convergence of the off-policy algorithm, and demonstrate that our approach enables the identification of policies that are finely tuned to the intricate risk-awareness of the agent that they serve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。