提出风险感知的安全强化学习,让智能体在噪声系统中既高效又安全。
Risk-Aware Safe Reinforcement Learning for Control of Stochastic Linear Systems
- 用可学习的双控制器结构替代传统干预,提升安全性。
- 通过优化单变量降低安全违规概率,减少对数据和模型的依赖。
- 适合需要高可靠性的自动化控制场景,如机器人、工业系统。
本文针对随机离散时间线性系统,提出一种风险感知的安全强化学习控制设计。不同于依赖高保真模型或即时干预的安全验证器,该方法同时学习一个风险感知的安全控制器与强化学习控制器,并通过数据驱动的插值融合二者。该框架具备三大优势:1)在有限数据下实现高置信度安全,无需精确系统模型;2)避免短视干预与收敛至不良平衡点,通过协调两个稳定控制器的贡献实现动态平衡;3)计算高效,仅需优化单一决策变量及线性规划多面体集。为获得更大不变集,采用分段仿射控制器而非线性控制器。基于采集数据,将闭环系统表示为决策变量与噪声的函数,形式化决策变量对安全违规方差的影响,并据此设计变量以最小化安全违规概率。结果表明该方法显著降低数据需求并减小安全违规方差。最后引入新型数据驱动插值技术,确保强化学习代理在噪声环境中保持最优性能的同时满足安全约束。仿真验证了理论成果的有效性。
原文摘要 · Abstract (English)
This paper presents a risk-aware safe reinforcement learning (RL) control design for stochastic discrete-time linear systems. Rather than using a safety certifier to myopically intervene with the RL controller, a risk-informed safe controller is also learned besides the RL controller, and the RL and safe controllers are combined together. Several advantages come along with this approach: 1) High-confidence safety can be certified without relying on a high-fidelity system model and using limited data available, 2) Myopic interventions and convergence to an undesired equilibrium can be avoided by deciding on the contribution of two stabilizing controllers, and 3) highly efficient and computationally tractable solutions can be provided by optimizing over a scalar decision variable and linear programming polyhedral sets. To learn safe controllers with a large invariant set, piecewise affine controllers are learned instead of linear controllers. To this end, the closed-loop system is first represented using collected data, a decision variable, and noise. The effect of the decision variable on the variance of the safe violation of the closed-loop system is formalized. The decision variable is then designed such that the probability of safety violation for the learned closed-loop system is minimized. It is shown that this control-oriented approach reduces the data requirements and can also reduce the variance of safety violations. Finally, to integrate the safe and RL controllers, a new data-driven interpolation technique is introduced. This method aims to maintain the RL agent's optimal implementation while ensuring its safety within environments characterized by noise. The study concludes with a simulation example that serves to validate the theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。