提出新算法优化风险敏感强化学习,可有效避开鞍点。
Policy Newton methods for Distortion Riskmetrics
- 基于似然比法推导风险目标的策略海森矩阵,设计采样估计器。
- 算法在样本复杂度 $\mathcal{O}(ε^{-3.5})$ 内收敛到 $ε$-二阶驻点。
- 首个实现风险敏感目标二阶收敛的算法,适合高风险决策场景。
我们研究强化学习框架下的风险敏感控制问题,旨在有限时域马尔可夫决策过程(MDP)中通过最大化折扣奖励的扭曲风险度量(DRM)来寻找风险最优策略。DRMs 是一类包含多种经典风险度量的丰富类。我们利用似然比方法推导了 DRM 目标关于策略的海森定理,并据此从 MDP 样本轨迹中提出了一个自然的 DRM 海森估计器。随后,我们在同策略强化学习设定下,提出了一种基于梯度与海森估计的立方正则化策略牛顿算法。该算法被证明能收敛至 DRM 目标的一个 $ε$-二阶驻点($ε$-SOSP),此保证可确保逃离鞍点。算法达到 $ε$-SOSP 的样本复杂度为 $\mathcal{O}(ε^{-3.5})$。实验验证了理论结果。据我们所知,这是首个针对风险敏感目标实现 $ε$-二阶驻点收敛的工作;此前文献要么仅保证风险敏感目标的一阶收敛,要么仅对风险无感目标实现二阶收敛。
原文摘要 · Abstract (English)
We consider the problem of risk-sensitive control in a reinforcement learning (RL) framework. In particular, we aim to find a risk-optimal policy by maximizing the distortion riskmetric (DRM) of the discounted reward in a finite horizon Markov decision process (MDP). DRMs are a rich class of risk measures that include several well-known risk measures as special cases. We derive a policy Hessian theorem for the DRM objective using the likelihood ratio method. Using this result, we propose a natural DRM Hessian estimator from sample trajectories of the underlying MDP. Next, we present a cubic-regularized policy Newton algorithm for solving this problem in an on-policy RL setting using estimates of the DRM gradient and Hessian. Our proposed algorithm is shown to converge to an $ε$-second-order stationary point ($ε$-SOSP) of the DRM objective, and this guarantee ensures the escaping of saddle points. The sample complexity of our algorithms to find an $ ε$-SOSP is $\mathcal{O}(ε^{-3.5})$. Our experiments validate the theoretical findings. To the best of our knowledge, our is the first work to present convergence to an $ε$-SOSP of a risk-sensitive objective, while existing works in the literature have either shown convergence to a first-order stationary point of a risk-sensitive objective, or a SOSP of a risk-neutral one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。