提出新方法实现约束强化学习的单策略收敛,解决理论与实践脱节问题。
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs

- 用增广拉格朗日法构建通用框架,确保单次策略迭代收敛
- 在表格和线性函数逼近下均实现全局收敛,无需混合策略
- 适用于复杂非线性策略,可在连续控制任务中有效应用
我们研究无限时域、折扣约束马尔可夫决策过程(CMDPs)的策略优化。现有理论通常仅对混合策略成立,但部署需单个策略,导致理论与实践不匹配。近期工作虽关注最后迭代收敛,但多限于表格设置或不实用的算法变体。为此,本文采用经典非精确增广拉格朗日(AL)方法,提出一个具有可证明最后迭代收敛性的通用框架。首先在表格设置下,使用投影Q上升(PQA)求解AL子问题;结合PQA的理论保证与标准AL分析,建立全局最后迭代收敛性。进一步推广至对数线性策略,证明一种高效的投影变体仍能获得与之前工作相当的收敛保证。最后,验证该框架可扩展至复杂非线性策略,并在连续控制任务中进行评估。
原文摘要 · Abstract (English)
We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture policy, deploying such a policy is computationally and memory intensive. This leads to a practical mismatch where a single (last-iterate) policy must be deployed. Recent theoretical works have thus focused on proving last-iterate convergence, but are largely limited to the tabular setting or to algorithmic variants that are rarely used in practice. To address this, we use the classic inexact augmented Lagrangian ($\texttt{AL}$) method from constrained optimization, and propose a general framework with provable last-iterate convergence for CMDPs. We first focus on the tabular setting and propose to solve the $\texttt{AL}$ sub-problem with projected Q-ascent ($\texttt{PQA}$). Combining the theoretical guarantees of $\texttt{PQA}$ and the standard $\texttt{AL}$ analysis enables us to establish global last-iterate convergence. We generalize these results to handle log-linear policies, and demonstrate that an efficient, projected variant of $\texttt{PQA}$ can achieve last-iterate convergence with comparable guarantees as prior work. Finally, we demonstrate that our framework scales to complex non-linear policies, and evaluate it on continuous control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。