arXiv:2410.07359eess.SYcs.LG2024-10被引 5

用数据驱动方法为未知动态的智能系统提供安全防护,确保不进入危险状态。

Learning-Based Shielding for Safe Autonomy under Unknown Dynamics

  • 基于深度核学习建模系统演化并量化不确定性
  • 构建区间MDP抽象,计算最大允许的安全策略集
  • 适用于黑箱控制器,如深度强化学习系统的安全验证

屏蔽是一种常用于保障黑箱控制器(如深度强化学习中的神经网络控制器)系统安全的方法,通过更简单的可验证控制器实现。现有屏蔽方法依赖于马尔可夫决策过程(MDP)的形式化验证,假设系统模型已知或为有限状态,这限制了其在未知、连续状态系统的应用。本文提出一种数据驱动的屏蔽方法,可在未知系统下保证安全。该方法利用深度核学习对系统一步演化进行建模,并量化不确定性,构建区间MDP(IMDP)抽象。针对以安全线性时序逻辑(safe LTL)表达的安全性质,设计算法求解最大化允许的安全策略集,确保避开不安全状态。理论证明了算法的正确性与计算复杂度,并在非线性系统,包括高维自主航天器场景中进行了实验验证。

原文摘要 · Abstract (English)

Shielding is a common method used to guarantee the safety of a system under a black-box controller, such as a neural network controller from deep reinforcement learning (DRL), with simpler, verified controllers. Existing shielding methods rely on formal verification through Markov Decision Processes (MDPs), assuming either known or finite-state models, which limits their applicability to DRL settings with unknown, continuous-state systems. This paper addresses these limitations by proposing a data-driven shielding methodology that guarantees safety for unknown systems under black-box controllers. The approach leverages Deep Kernel Learning to model the systems' one-step evolution with uncertainty quantification and constructs a finite-state abstraction as an Interval MDP (IMDP). By focusing on safety properties expressed in safe linear temporal logic (safe LTL), we develop an algorithm that computes the maximally permissive set of safe policies on the IMDP, ensuring avoidance of unsafe states. The algorithms soundness and computational complexity are demonstrated through theoretical proofs and experiments on nonlinear systems, including a high-dimensional autonomous spacecraft scenario.

安全控制强化学习形式验证数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。