arXiv:2502.10138cs.LG2025-02NeurIPS被引 3

提出高效安全强化学习算法,每轮不违规且误差随轮数平方根增长。

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

  • 基于线性函数近似设计新算法,保证每轮不违反约束。
  • 理论证明误差上界为$ ilde{ m O}( ext{√}K)$,优于已有方法。
  • 适合需严格安全约束的复杂环境决策任务。

研究在受限马尔可夫决策过程(CMDP)中的强化学习问题,要求智能体在最大化累积奖励的同时,确保每轮中期望总效用值满足单一约束。尽管该问题在表格情形下已清晰,但函数逼近下的理论结果仍匮乏。本文填补空白,提出一种适用于线性CMDP的强化学习算法,实现$ ilde{ m O}( ext{√}K)$的遗憾上界,并保证每轮零违规。此外,该方法计算效率高,复杂度多项式依赖于问题相关参数,独立于状态空间大小。相比近期线性CMDP算法,本方法既避免约束违规,又无指数级计算开销。

原文摘要 · Abstract (English)

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well understood in the tabular setting, theoretical results for function approximation remain scarce. This paper closes the gap by proposing an RL algorithm for linear CMDPs that achieves $\tilde{\mathcal{O}}(\sqrt{K})$ regret with an episode-wise zero-violation guarantee. Furthermore, our method is computationally efficient, scaling polynomially with problem-dependent parameters while remaining independent of the state space size. Our results significantly improve upon recent linear CMDP algorithms, which either violate the constraint or incur exponential computational costs.

强化学习安全控制线性函数逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。