提出可扩展的安全强化学习方法,确保智能体训练测试全程安全。
Probabilistic Shielding for Safe Reinforcement Learning
- 通过状态扩展构建安全屏障,动态限制智能体动作选择
- 在已知安全动态条件下,实现严格形式化安全保证
- 适合对安全性要求高、需实时保障的工业级强化学习应用
在真实场景中,强化学习(RL)智能体在追求最大奖励的同时,必须在训练阶段也保持安全行为。近年来,安全强化学习受到广泛关注,其目标是在满足给定安全约束的所有策略中学习最优策略。然而,现有方法多基于线性规划,难以扩展。本文提出一种新方法,在已知马尔可夫决策过程(MDP)安全动态且安全定义为无折扣概率规避属性的前提下,实现可扩展且具有严格形式化安全保证的强约束安全强化学习。该方法基于对MDP的状态扩展,并设计一个限制智能体可用动作的防护盾。我们证明该方法能严格保证智能体在训练和测试阶段均保持安全。此外,实验验证了该方法在实际中的可行性。
原文摘要 · Abstract (English)
In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximise their reward, must often also behave in a safe manner, including at training time. Thus, much attention in recent years has been given to Safe RL, where an agent aims to learn an optimal policy among all policies that satisfy a given safety constraint. However, strict safety guarantees are often provided through approaches based on linear programming, and thus have limited scaling. In this paper we present a new, scalable method, which enjoys strict formal guarantees for Safe RL, in the case where the safety dynamics of the Markov Decision Process (MDP) are known, and safety is defined as an undiscounted probabilistic avoidance property. Our approach is based on state-augmentation of the MDP, and on the design of a shield that restricts the actions available to the agent. We show that our approach provides a strict formal safety guarantee that the agent stays safe at training and test time. Furthermore, we demonstrate that our approach is viable in practice through experimental evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。