首个在安全约束下实现√T悔值的在线强化学习算法
Foundations of Safe Online Reinforcement Learning in the Linear Quadratic Regulator: $\sqrt{T}$-Regret
- 设计安全约束下的线性系统控制算法,保证状态高概率在安全区间
- 首次达成~O(√T)悔值,优于无约束情形的收敛速度
- 适用于需保障安全的工业控制场景,如机器人与自动驾驶
理解如何在遵守安全约束的前提下高效学习,对在线强化学习的实际应用至关重要。然而,由于安全、探索与利用之间的复杂交互,证明安全约束强化学习的严格悔值界十分困难。本文通过研究一维未知动态线性系统的经典控制问题,建立安全约束强化学习的基础。在状态需以高概率保持在安全区域的约束下,提出首个达到~O_T(√T)悔值的安全算法。该悔值基准为截断线性控制器,是适用于安全约束线性系统的自然非线性控制器基准。此外,我们还证明了该基准下最优控制器的若干良好连续性性质。在主结果证明中发现,当约束影响最优控制器时,所用控制器类的非线性特性导致学习速率快于无约束情形。
原文摘要 · Abstract (English)
Understanding how to efficiently learn while adhering to safety constraints is essential for using online reinforcement learning in practical applications. However, proving rigorous regret bounds for safety-constrained reinforcement learning is difficult due to the complex interaction between safety, exploration, and exploitation. In this work, we seek to establish foundations for safety-constrained reinforcement learning by studying the canonical problem of controlling a one-dimensional linear dynamical system with unknown dynamics. We study the safety-constrained version of this problem, where the state must with high probability stay within a safe region, and we provide the first safe algorithm that achieves regret of $\tilde{O}_T(\sqrt{T})$. Furthermore, the regret is with respect to the baseline of truncated linear controllers, a natural baseline of non-linear controllers that are well-suited for safety-constrained linear systems. In addition to introducing this new baseline, we also prove several desirable continuity properties of the optimal controller in this baseline. In showing our main result, we prove that whenever the constraints impact the optimal controller, the non-linearity of our controller class leads to a faster rate of learning than in the unconstrained setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。