arXiv:2410.22114cs.LGcs.AI2024-10被引 5

提出可全局最优的鲁棒强化学习策略梯度方法

Policy Gradient for Robust Markov Decision Processes

  • 采用双循环镜面下降优化,自适应调整容错率
  • 在多种设置下实现策略全局收敛与鲁棒性验证
  • 适合需应对不确定性模型的强化学习应用

我们为鲁棒马尔可夫决策过程(Robust MDPs)开发了一种具备全局最优保证的通用策略梯度方法。尽管策略梯度方法因可扩展性和高效性被广泛用于动态决策问题,但将其应用于模型不确定性的场景仍具挑战,常难以学习到鲁棒策略。本文提出一种新型策略梯度方法——双循环鲁棒策略镜面下降(DRPMD),采用通用镜面下降更新规则,并在每轮迭代中自适应调整容忍度,确保收敛至全局最优策略。我们对DRPMD进行了全面分析,包括直接参数化与Softmax参数化下的新收敛结果,并通过转移镜面上升(TMA)揭示了内层问题求解的新见解。此外,我们提出了适用于离散与连续状态-动作空间的创新参数化转移核,拓展了方法的应用范围。实验结果验证了DRPMD在多种复杂鲁棒MDP场景下的鲁棒性与全局收敛性。

原文摘要 · Abstract (English)

We develop a generic policy gradient method with the global optimality guarantee for robust Markov Decision Processes (MDPs). While policy gradient methods are widely used for solving dynamic decision problems due to their scalable and efficient nature, adapting these methods to account for model ambiguity has been challenging, often making it impractical to learn robust policies. This paper introduces a novel policy gradient method, Double-Loop Robust Policy Mirror Descent (DRPMD), for solving robust MDPs. DRPMD employs a general mirror descent update rule for the policy optimization with adaptive tolerance per iteration, guaranteeing convergence to a globally optimal policy. We provide a comprehensive analysis of DRPMD, including new convergence results under both direct and softmax parameterizations, and provide novel insights into the inner problem solution through Transition Mirror Ascent (TMA). Additionally, we propose innovative parametric transition kernels for both discrete and continuous state-action spaces, broadening the applicability of our approach. Empirical results validate the robustness and global convergence of DRPMD across various challenging robust MDP settings.

强化学习鲁棒优化策略梯度马尔可夫决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。