提出一种用于风险敏感成本的函数逼近强化学习算法。
An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes
- 基于演员-评论家框架,结合函数逼近处理风险敏感成本问题。
- 理论证明算法渐近收敛,数值实验显示性能优于现有方法。
- 适合关注风险控制与复杂决策场景的研究者。
本文研究马尔可夫决策过程中的风险敏感成本准则,采用指数化成本形式,并在此设定下开发了一种无模型策略梯度算法。与平均或折扣成本等加性准则不同,风险敏感成本准则因导致贝尔曼方程具有乘法结构而更难处理。本文提出一种带函数逼近的演员-评论家算法,并提供了其渐近收敛性分析。此外,数值实验结果表明,该算法在性能上优于文献中其他近期算法。
原文摘要 · Abstract (English)
In this paper, we consider the risk-sensitive cost criterion with exponentiated costs for Markov decision processes and develop a model-free policy gradient algorithm in this setting. Unlike additive cost criteria such as average or discounted cost, the risk-sensitive cost criterion is less studied due to the complexity resulting from the multiplicative structure of the resulting Bellman equation. We develop an actor-critic algorithm with function approximation in this setting and provide its asymptotic convergence analysis. We also show the results of numerical experiments that demonstrate the superiority in performance of our algorithm over other recent algorithms in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。