arXiv:2602.10917cs.LG2026-02

首个实现近恒定约束违规的在线强化学习算法

Near-Constant Strong Violation and Last-Iterate Convergence for Online CMDPs via Decaying Safety Margins

  • 引入动态安全边界与正则化,重构原始-对偶框架
  • 强约束违规维持在近常数水平($ ilde{O}(1)$),同时强收益遗憾亚线性
  • 支持非渐近最后迭代收敛,适合严格安全要求场景

我们研究在强遗憾和强违规度量下,约束马尔可夫决策过程(CMDPs)中的安全在线强化学习问题,该度量禁止时间上的误差抵消。现有原始-对偶方法虽能实现亚线性强收益遗憾,但不可避免导致累积强约束违规持续增长,或受限于平均迭代收敛,因内在振荡所致。为此,我们提出灵活安全域优化通过边际正则化探索(FlexDOME)算法,首次在理论上证明可实现近恒定 $ ilde{O}(1)$ 的强约束违规,同时保持亚线性强遗憾和非渐近最后迭代收敛。FlexDOME 在原始-对偶框架中引入时变安全边界与正则化项。理论分析基于一种新颖的逐项渐近主导策略:安全边界被严格调度,以渐近主导优化与统计误差的函数衰减速率,从而将累积违规钳制在近恒定水平。此外,通过策略-对偶李雅普诺夫论证,建立了非渐近最后迭代收敛保证。实验验证了理论结果。

原文摘要 · Abstract (English)

We study safe online reinforcement learning in Constrained Markov Decision Processes (CMDPs) under strong regret and violation metrics, which forbid error cancellation over time. Existing primal-dual methods that achieve sublinear strong reward regret inevitably incur growing strong constraint violation or are restricted to average-iterate convergence due to inherent oscillations. To address these limitations, we propose the Flexible safety Domain Optimization via Margin-regularized Exploration (FlexDOME) algorithm, the first to provably achieve near-constant $\tilde{O}(1)$ strong constraint violation alongside sublinear strong regret and non-asymptotic last-iterate convergence. FlexDOME incorporates time-varying safety margins and regularization terms into the primal-dual framework. Our theoretical analysis relies on a novel term-wise asymptotic dominance strategy, where the safety margin is rigorously scheduled to asymptotically majorize the functional decay rates of the optimization and statistical errors, thereby clamping cumulative violations to a near-constant level. Furthermore, we establish non-asymptotic last-iterate convergence guarantees via a policy-dual Lyapunov argument. Experiments corroborate our theoretical findings.

强化学习在线学习约束优化收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。