arXiv:2410.21081stat.MLcs.LG2024-10

提出安全约束下强化学习的新框架,实现更优的在线学习性能。

Foundations of Safe Online Reinforcement Learning in the Linear Quadratic Regulator: Generalized Baselines

  • 构建非线性控制器的通用分析框架,适配安全约束场景。
  • 在噪声支持集充足时,达到 ilde{O}_T(\ ext{\sqrt{T}})的后悔界。
  • 揭示噪声可带来免费探索,缓解安全约束带来的不确定性成本。

许多在线强化学习的实际应用需要在学习未知环境的同时满足安全约束。本文通过研究线性二次调节器(LQR)中动态未知但状态需全程高概率处于安全区域的经典问题,建立了安全强化学习的理论基础。主要贡献是提出一个适用于非线性控制器的通用框架,其在约束问题中表现优于传统线性控制器。由于非线性控制器在约束条件下分析困难,研究聚焦于一维状态与动作空间,但也讨论了高维推广的可能性。基于该框架,证明了在任意满足自然假设的非线性基线条件下,当噪声分布具有足够大的支撑集时,可实现 ilde{O}_T( ext{\sqrt{T}})的后悔率;对任意亚高斯噪声,则可实现 ilde{O}_T(T^{2/3})的后悔率。推导过程中引入一种新的非线性控制不确定性估计边界,表明在充分噪声条件下,安全约束反而能提供‘免费探索’,补偿安全约束带来的额外不确定性代价。

原文摘要 · Abstract (English)

Many practical applications of online reinforcement learning require the satisfaction of safety constraints while learning about the unknown environment. In this work, we establish theoretical foundations for reinforcement learning with safety constraints by studying the canonical problem of Linear Quadratic Regulator learning with unknown dynamics, but with the additional constraint that the position must stay within a safe region for the entire trajectory with high probability. Our primary contribution is a general framework for studying stronger baselines of nonlinear controllers that are better suited for constrained problems than linear controllers. Due to the difficulty of analyzing non-linear controllers in a constrained problem, we focus on 1-dimensional state- and action- spaces, however we also discuss how we expect the high-level takeaways can generalize to higher dimensions. Using our framework, we show that for \emph{any} non-linear baseline satisfying natural assumptions, $\tilde{O}_T(\sqrt{T})$-regret is possible when the noise distribution has sufficiently large support, and $\tilde{O}_T(T^{2/3})$-regret is possible for \emph{any} subgaussian noise distribution. In proving these results, we introduce a new uncertainty estimation bound for nonlinear controls which shows that enforcing safety in the presence of sufficient noise can provide ``free exploration'' that compensates for the added cost of uncertainty in safety-constrained control.

强化学习安全控制在线学习后悔界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。