兼顾安全与性能的强化学习控制方法,可应对未知扰动和故障。
Safe Learning Control with Optimality and Stability Guarantees
- 提出高阶倒数型控制屏障函数,处理高相对阶约束。
- 无需扰动上界信息,仍能保证系统鲁棒安全。
- 通过梯度调节提升性能,适合复杂工业控制场景。
单纯追求性能可能损害安全性,而过于保守的安全探索又会降低性能。如何在基于学习的控制问题中同时保障安全与性能,是一个重要且具有挑战性的问题。本文针对存在高相对阶状态约束及未知时变扰动/执行器故障的非线性系统,通过求解基于强化学习(RL)的最优控制问题,提升系统性能的同时确保安全。提出一种新型控制屏障函数(CBF),称为高阶倒数型控制屏障函数,可处理高相对阶约束,并在不依赖扰动上界的情况下实现鲁棒安全。引入梯度相似性概念,量化安全与性能之间的关系。在基于模型的安全强化学习框架中,结合梯度调节与自适应机制,在保证安全的前提下提升性能。两个仿真案例验证了所提算法的有效性。
原文摘要 · Abstract (English)
Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance. How to guarantee both safety and performance in learning-based control problems is an interesting yet challenging issue. This paper aims to enhance system performance with a safety guarantee by solving reinforcement learning (RL)-based optimal control problems for nonlinear systems subject to high-relative-degree state constraints and unknown time-varying disturbance/actuator faults. A new type of control barrier functions (CBFs), termed high-order reciprocal-based control barrier function, is proposed to handle high-relative-degree constraints, which extends the design of CBFs to enforce robust safety without knowing the disturbance bound. The concept of gradient similarity is proposed to quantify the relationship between safety and performance. Finally, gradient manipulation and adaptive mechanisms are introduced in the model-based safe RL framework to enhance the performance with a safety guarantee. Two simulation examples illustrate the efficacy of the proposed algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。