为概率安全设计新防护机制,兼顾安全性与灵活性。
Shields to Guarantee Probabilistic Safety in MDPs
- 提出可保守扩展的经典防护框架,支持允许低概率风险。
- 证明无法同时保持强安全与高放行性,需权衡取舍。
- 提供离线与在线构建方法,确保强安全且计算可行。
防护(Shielding)是保障自主智能体安全的主流模型驱动技术。经典防护保证永不发生坏事,具有强安全性和最大放行性。然而,针对允许以可接受概率发生坏事的概率安全场景,防护系统的设计更为复杂。本文提出一个形式化框架,保守扩展经典防护至概率安全情形。在该框架中,(i) 证明了无法同时维持强安全性和最大放行性;(ii) 提供具有较弱保证的自然防护;(iii) 引入离线与在线防护构造方法,实现强安全保证。实证评估表明,新防护机制兼具实际优势与计算可行性。
原文摘要 · Abstract (English)
Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens and comes with strong guarantees about safety and maximal permissiveness. However, shielding systems for probabilistic safety, where something bad is allowed to happen with an acceptable probability, has proven to be more intricate. This paper presents a formal framework that conservatively extends classical shields to probabilistic safety. In this framework, we (i) demonstrate the impossibility of preserving the strong guarantees on safety and permissiveness, (ii) provide natural shields with weaker guarantees, and (iii) introduce offline and online shield constructions ensuring strong safety guarantees. The empirical evaluation highlights the practical advantages of the new shields, as well as their computational feasibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。