提出三类风险敏感强化学习算法,提升决策鲁棒性。
Risk-sensitive reinforcement learning using expectiles, shortfall risk and optimized certainty equivalent risk
- 基于期望值、短缺风险等三类风险度量设计策略梯度方法
- 理论证明估计误差为1/m,收敛速度有保障
- 适合高风险场景下的安全强化学习研究者
我们提出了针对三类风险度量(期望值、基于效用的短缺风险、优化确定等价风险)的风险敏感强化学习算法。在有限时域马尔可夫决策过程框架下,首先推导了相应的策略梯度定理;其次,为每类风险度量设计了风险敏感策略梯度估计器,并建立了均方误差为 $\mathcal{O}(1/m)$ 的理论边界,其中 $m$ 为轨迹数量;在标准策略梯度假设下,进一步证明了风险敏感目标函数的光滑性,从而获得所提算法的整体收敛速率上界;最后,在主流强化学习基准上进行了数值实验,验证了理论结果的有效性。
原文摘要 · Abstract (English)
We propose risk-sensitive reinforcement learning algorithms catering to three families of risk measures, namely expectiles, utility-based shortfall risk and optimized certainty equivalent risk. For each risk measure, in the context of a finite horizon Markov decision process, we first derive a policy gradient theorem. Second, we propose estimators of the risk-sensitive policy gradient for each of the aforementioned risk measures, and establish $\mathcal{O}\left(1/m\right)$ mean-squared error bounds for our estimators, where $m$ is the number of trajectories. Further, under standard assumptions for policy gradient-type algorithms, we establish smoothness of the risk-sensitive objective, in turn leading to stationary convergence rate bounds for the overall risk-sensitive policy gradient algorithm that we propose. Finally, we conduct numerical experiments to validate the theoretical findings on popular RL benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。