arXiv:2607.28390cs.LG2026-07

提出新型神经批判者算法,实现安全强化学习的最优收敛速度。

Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

  • 用分层多级蒙特卡洛技术同时降低采样与优化偏差。
  • 在平均奖励约束下实现 $ ilde{O}(T^{-1/2})$ 的最优误差与约束违反率。
  • 无需知道系统混合时间,适合复杂安全场景中的深度强化学习。

约束马尔可夫决策过程(CMDPs)为安全关键型强化学习提供了自然框架,要求智能体在最大化长期回报的同时满足长期约束。尽管基于线性批判者的原始-对偶演员-批判者方法已有充分理解,但在平均奖励CMDPs中将最优收敛性保证扩展到神经批判者仍是一个开放问题。主要挑战在于神经批判者估计中的偏差-成本权衡:在神经正切核(NTK)分析下,显著降低偏差会大幅增加优化成本,导致原始-对偶框架无法实现最优收敛。本文通过引入分层多级蒙特卡洛(MLMC)神经批判者,实现了轨迹采样与批判者优化过程中的同步去偏。该估计器以对数期望样本成本达到长时优化运行的偏差水平。基于此,我们构建了原始-对偶自然演员-批判者算法,实现了最优差距和约束违反度均为 $ ilde{O}(T^{-1/2})$。这首次在一般策略参数化和神经批判者条件下,建立了无限时域平均奖励CMDPs的最优收敛性保证,且无需预知底层混合时间。结果在无约束情形下亦具新颖性。

原文摘要 · Abstract (English)

Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maximize long-term reward while satisfying long-term constraints. Although primal-dual actor-critic methods with linear critics are well understood, extending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained open. The main challenge is a fundamental bias-cost trade-off in neural critic estimation: under Neural Tangent Kernel (NTK) analysis, reducing critic bias substantially increases critic optimization cost, preventing order-optimal convergence in the primal-dual framework. We resolve this bottleneck by introducing a hierarchical Multilevel Monte Carlo (MLMC) neural critic that performs debiasing simultaneously across trajectory sampling and critic optimization. The resulting estimator attains the bias of a long critic optimization run with only logarithmic expected sample cost. Building on this estimator, we develop a primal-dual Natural Actor-Critic algorithm that achieves both an optimality gap and a constraint violation of order $\tilde{O}(T^{-1/2})$. This establishes the first order-optimal convergence guarantees for infinite-horizon average-reward CMDPs with general policy parameterization and neural critics, while eliminating the need to know the underlying mixing time. Our results are novel even in the unconstrained setting.

强化学习神经网络约束优化收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。