arXiv:2412.12996cs.LGcs.AI2024-12AAAI被引 6

用运行时监控修复神经网络控制策略的安全缺陷

Neural Control and Certificate Repair via Runtime Monitoring

  • 通过运行时监测发现违反安全属性的行为,生成新训练数据
  • 在两个自主系统任务中将安全率显著提升,修复了学习到的控制策略
  • 适合关注神经控制安全性的研究者和工业应用开发者

基于学习的方法为解决高度非线性控制任务提供了可行路径,传统方法难以应对。为确保安全属性满足,这类方法联合学习控制策略与证明该属性成立的证书函数,如用于安全性的屏障函数、用于渐近稳定的李雅普诺夫函数。尽管在白盒环境下(系统动态已知)已有显著进展,但在黑盒环境(系统动态未知)下如何保障证书函数可靠性仍缺乏研究。本文提出一种新框架,利用运行时监控检测初始神经网络策略与证书函数下的违规行为,并据此提取新训练数据,重新训练策略与证书以实现修复。我们在两个自主系统控制任务上实证验证了该方法的有效性,成功提升了基于先进学习控制方法所学策略的安全率。

原文摘要 · Abstract (English)

Learning-based methods provide a promising approach to solving highly non-linear control tasks that are often challenging for classical control methods. To ensure the satisfaction of a safety property, learning-based methods jointly learn a control policy together with a certificate function for the property. Popular examples include barrier functions for safety and Lyapunov functions for asymptotic stability. While there has been significant progress on learning-based control with certificate functions in the white-box setting, where the correctness of the certificate function can be formally verified, there has been little work on ensuring their reliability in the black-box setting where the system dynamics are unknown. In this work, we consider the problems of certifying and repairing neural network control policies and certificate functions in the black-box setting. We propose a novel framework that utilizes runtime monitoring to detect system behaviors that violate the property of interest under some initially trained neural network policy and certificate. These violating behaviors are used to extract new training data, that is used to re-train the neural network policy and the certificate function and to ultimately repair them. We demonstrate the effectiveness of our approach empirically by using it to repair and to boost the safety rate of neural network policies learned by a state-of-the-art method for learning-based control on two autonomous system control tasks.

神经控制安全验证运行时监控证书修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。