首个证明神经网络批评家在约束强化学习中全局收敛的理论工作
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
- 用自然策略梯度与神经网络批评家结合的原始对偶算法
- 收敛速度达T⁻¹/⁴,约束违反累积率与函数逼近误差相关
- 适用于高维连续控制,突破线性批评家限制
研究无限时域约束马尔可夫决策过程(CMDPs)在通用策略参数化和多层神经网络批评家下的理论性质。现有约束强化学习理论多依赖表格策略或线性批评家,难以应用于高维和连续控制问题。本文提出一种原始对偶自然演员-评论家算法,将神经网络批评家估计与自然策略梯度更新结合,并利用神经正切核(NTK)理论,在非平稳采样下控制函数逼近误差,无需混合时间预言机。建立了全局收敛性和累积约束违反率 $ ilde{ ext{O}}(T^{-1/4})$ 的理论保证,该结果首次适用于具有通用策略和多层神经批评家的CMDPs,显著拓展了演员-评论家方法的理论基础。
原文摘要 · Abstract (English)
We study infinite-horizon Constrained Markov Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on tabular policies or linear critics, which limits their applicability to high-dimensional and continuous control problems. We propose a primal-dual natural actor-critic algorithm that integrates neural critic estimation with natural policy gradient updates and leverages Neural Tangent Kernel (NTK) theory to control function-approximation error under Markovian sampling, without requiring access to mixing-time oracles. We establish global convergence and cumulative constraint violation rates of $\tilde{\mathcal{O}}(T^-1/4)$ up to approximation errors induced by the policy and critic classes. Our results provide the first such guarantees for CMDPs with general policies and multi-layer neural critics, substantially extending the theoretical foundations of actor-critic methods beyond the linear-critic regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。