用系统理论分析RL安全风险,发现传统方法遗漏的隐患。
RL-STPA: Adapting System-Theoretic Hazard Analysis for Safety-Critical Reinforcement Learning
- 通过任务分解和专家知识识别潜在失控行为
- 用扰动测试覆盖状态动作空间,发现隐蔽风险点
- 可迭代优化训练,适合无人机等高危场景
随着强化学习在安全关键领域应用扩展,现有评估方法难以系统识别由神经网络策略黑箱特性和训练部署分布偏移带来的安全隐患。本文提出强化学习系统理论过程分析(RL-STPA),通过三项核心贡献将传统STPA方法适配至RL场景:利用时序阶段分析与领域知识进行分层子任务分解以捕捉涌现行为;设计覆盖引导的扰动测试,探索状态-动作空间敏感性;通过迭代检查点将识别出的风险反馈至训练,实现奖励重塑与课程设计。在自主无人机导航与着陆这一安全关键案例中,该框架揭示了标准强化学习评估可能遗漏的损失场景。所提框架为从业者提供系统化危害分析工具、安全性覆盖度量化指标及操作安全边界制定指南。尽管无法对任意神经策略提供形式化保证,但在难以实施穷尽验证的安全关键应用中,提供了实用的强化学习安全性与鲁棒性评估与改进方法。
原文摘要 · Abstract (English)
As reinforcement learning (RL) deployments expand into safety-critical domains, existing evaluation methods fail to systematically identify hazards arising from the black-box nature of neural network enabled policies and distributional shift between training and deployment. This paper introduces Reinforcement Learning System-Theoretic Process Analysis (RL-STPA), a framework that adapts conventional STPA's systematic hazard analysis to address RL's unique challenges through three key contributions: hierarchical subtask decomposition using both temporal phase analysis and domain expertise to capture emergent behaviors, coverage-guided perturbation testing that explores the sensitivity of state-action spaces, and iterative checkpoints that feed identified hazards back into training through reward shaping and curriculum design. We demonstrate RL-STPA in the safety-critical test case of autonomous drone navigation and landing, revealing potential loss scenarios that can be missed by standard RL evaluations. The proposed framework provides practitioners with a toolkit for systematic hazard analysis, quantitative metrics for safety coverage assessment, and actionable guidelines for establishing operational safety bounds. While RL-STPA cannot provide formal guarantees for arbitrary neural policies, it offers a practical methodology for systematically evaluating and improving RL safety and robustness in safety-critical applications where exhaustive verification methods remain intractable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。