提出风险感知的通用效用马尔可夫决策过程,实现安全与性能的权衡。
Risk-Aware General-Utility Markov Decision Processes

- 基于熵风险度量设计风险感知策略,平衡期望收益与风险厌恶。
- 采用蒙特卡洛树搜索实现任意精度求解,理论保证可行。
- 在探索、模仿学习等多场景中验证有效,适合高风险决策任务。
我们研究具有风险感知目标的通用效用马尔可夫决策过程(GUMDPs)。在此框架中,智能体旨在优化目标值分布的风险度量,而目标函数依赖于策略诱导的状态访问频率。首先,我们提出并形式化了风险感知的GUMDPs,使决策者可在期望性能与风险规避间权衡,并利用该框架支持丰富的目标形式。重点关注熵风险度量(ERM)。其次,通过在线规划技术求解带有ERM目标的风险感知GUMDPs,提出一种基于蒙特卡洛树搜索(MCTS)的方法,可保证在任意精度下求解。第三,实验表明,该方法在标准MDP、最大状态熵探索、模仿学习及多目标MDP等多种任务中,成功优化了多种风险感知行为。
原文摘要 · Abstract (English)
We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk aversion while benefiting from the rich set of objectives that can be cast under the framework of GUMDPs. We focus our attention on the entropic risk measure (ERM). Second, we show how we can solve risk-aware GUMDPs with ERM objectives by resorting to online planning techniques. In particular, we propose an approach based on Monte Carlo Tree Search (MCTS) to provably solve risk-aware GUMDPs up to any desired accuracy. Third, we provide a set of experimental results showcasing that our approach is successful when optimizing for a spectrum of risk-aware behaviors in the context of GUMDPs under diverse tasks (standard MDPs, maximum state entropy exploration, imitation learning, and multi-objective MDPs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。