针对不确定参数的强化学习,提出贝叶斯风险敏感策略优化方法
Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
- 用贝叶斯方法估计未知参数,并在后验分布上施加一致风险度量
- 提出新梯度算法,收敛速度达 $\mathcal{O}(T^{-1/2}+r^{-1/2})$
- 适用于高风险敏感场景,如金融、医疗决策等
针对具有通用损失函数和未知参数的马尔可夫决策过程(MDPs),为缓解由参数不确定性带来的认知不确定性,采用贝叶斯方法从数据中估计参数,并在贝叶斯后验分布上施加一致风险泛函。由于该形式通常不满足交换性原则,无法使用基于动态规划的贝尔曼方程求解。为此,本文提出一种策略梯度优化方法,利用一致风险度量的对偶表示,并将包络定理推广至连续情形。我们证明了该算法的平稳性分析,收敛速率为 $\mathcal{O}(T^{-1/2}+r^{-1/2})$,其中 $T$ 为策略梯度迭代次数,$r$ 为梯度估计器的样本量。进一步将算法扩展到回合制设置,建立了全局收敛性,并给出了每回合达到误差界 $\mathcal{O}(ε)$ 所需的迭代次数上界。
原文摘要 · Abstract (English)
Motivated by many application problems, we consider Markov decision processes (MDPs) with a general loss function and unknown parameters. To mitigate the epistemic uncertainty associated with unknown parameters, we take a Bayesian approach to estimate the parameters from data and impose a coherent risk functional (with respect to the Bayesian posterior distribution) on the loss. Since this formulation usually does not satisfy the interchangeability principle, it does not admit Bellman equations and cannot be solved by approaches based on dynamic programming. Therefore, We propose a policy gradient optimization method, leveraging the dual representation of coherent risk measures and extending the envelope theorem to continuous cases. We then show the stationary analysis of the algorithm with a convergence rate of $\mathcal{O}(T^{-1/2}+r^{-1/2})$, where $T$ is the number of policy gradient iterations and $r$ is the sample size of the gradient estimator. We further extend our algorithm to an episodic setting, and establish the global convergence of the extended algorithm and provide bounds on the number of iterations needed to achieve an error bound $\mathcal{O}(ε)$ in each episode.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。