为强化学习系统提供可证明的安全保障框架,解决安全关键场景下的可靠性问题。
Assuring the Safety of Reinforcement Learning Components: AMLAS-RL
- 基于AMLAS方法迭代构建强化学习系统的安全论证
- 在轮式车辆避障任务中验证了框架的有效性
- 适合安全关键型智能系统开发者使用
机器学习的快速发展使其在多个领域的网络物理系统(CPS)中得到广泛应用。尽管CPS具备强大能力,但引入机器学习组件带来了显著的安全与可保证性挑战。强化学习(RL)因其在复杂动态环境中无需显式建模即可学习的优势,特别适用于CPS。然而,在安全关键应用中,学习过程不仅需高效,更需可证明安全。现有安全强化学习方法虽在训练中加入安全约束,但未能系统性覆盖整个RL生命周期。现有的AMLAS方法为监督学习组件提供了结构化安全保障指导,但不直接适用于强化学习的独特挑战。本文提出将AMLAS适配为AMLAS-RL框架,通过迭代过程生成强化学习系统的确证论据。以一辆轮式车辆在无碰撞前提下到达目标点的任务为例,展示了该方法的可行性与实用性。
原文摘要 · Abstract (English)
The rapid advancement of machine learning (ML) has led to its increasing integration into cyber-physical systems (CPS) across diverse domains. While CPS offer powerful capabilities, incorporating ML components introduces significant safety and assurance challenges. Among ML techniques, reinforcement learning (RL) is particularly suited for CPS due to its capacity to handle complex, dynamic environments where explicit models of interaction between system and environment are unavailable or difficult to construct. However, in safety-critical applications, this learning process must not only be effective but demonstrably safe. Safe-RL methods aim to address this by incorporating safety constraints during learning, yet they fall short in providing systematic assurance across the RL lifecycle. The AMLAS methodology offers structured guidance for assuring the safety of supervised learning components, but it does not directly apply to the unique challenges posed by RL. In this paper, we adapt AMLAS to provide a framework for generating assurance arguments for an RL-enabled system through an iterative process; AMLAS-RL. We demonstrate AMLAS-RL using a running example of a wheeled vehicle tasked with reaching a target goal without collision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。