提出新算法实现约束LQR的最优后悔率,保障每步安全。
Rate-Optimal Regret for the Safe Learning-based Control of the Constrained Linear Quadratic Regulator
- 用半定规划选乐观策略,再缩放确保安全
- 达到理论最优的√T后悔率,满足概率约束
- 适合关注安全强化学习与在线控制的研究者
研究带有约束的随机线性二次调节器(LQR)自适应控制问题,要求每一步都满足约束。已有工作在多维情形下实现了˜O(T^{2/3})的后悔率并保证鲁棒约束,但未解决能否在约束下达到˜O(√T)后悔率的问题。本文证明了在概率约束下可实现˜O(√T)后悔率,该类约束允许处理无界噪声,并支持非直接适用于鲁棒约束的分析技术。所提算法通过半定规划(SDP)选择乐观策略,再逐步缩放直至验证安全。理论分析基于一个关键引理:将系统协方差与所选策略关联,实现对后悔和约束的双重保证。该协方差分析方法区别于传统基于代价-到-目标的分析框架。
原文摘要 · Abstract (English)
We study the problem of adaptive control of the stochastic linear quadratic regulator (LQR) with constraints that must be satisfied at every time step. Prior work on the multidimensional problem has shown $\tilde{O}(T^{2/3})$ regret and satisfaction of robust constraints, leaving open the question of whether $\tilde{O}(\sqrt{T})$ regret can be attained in the constrained LQR setting. We contribute to this problem by showing $\tilde{O}(\sqrt{T})$ regret and satisfaction of chance constraints. This type of constraints allow us to handle unbounded noise and also enable analytical techniques not directly applicable to robust constraints. Our proposed algorithm for this problem uses an SDP to select an optimistic policy, and then "scales back" this policy until it is verifiably-safe. Our theoretical analysis establishes regret and constraint guarantees via a key lemma that bounds the system covariance in terms of the chosen policy. This covariance-based analysis is in contrast with the cost-to-go based analysis that is typically used in adaptive LQR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。