提出安全策略比方法,实现无模型强化学习中任务策略的安全训练与部署。
SPoRt -- Safe Policy Ratio: Certified Training and Deployment of Task Policies in Model-Free RL
- 基于安全基策略计算最大策略比,给出新策略违反安全属性的概率上界。
- 实验验证了安全与性能的可调和性,理论边界与实测结果吻合良好。
- 适合关注强化学习安全性、需可证明保障的工业级应用开发者。
为将强化学习应用于安全关键场景,必须在策略训练与部署阶段提供安全保证。本文提出了理论结果,给出了在无模型、周期性设置下,针对新任务策略违反安全属性的概率上界。该上界基于相对于一个‘安全’基策略计算的最大策略比,且可扩展至时序延展性质(如长期安全)与鲁棒控制问题。为应用该理论,我们提出SPoRt,其采用场景法实现对基策略的上界数据驱动计算,并引入投影PPO(Projected PPO),一种基于投影的训练方法,在保证用户指定的属性违规概率上界前提下训练任务特定策略。因此,SPoRt允许用户在安全保证与任务性能之间进行权衡。此外,实验展示了该权衡关系,并将理论边界与基于经验违规率的后验边界进行了对比,验证了其有效性。
原文摘要 · Abstract (English)
To apply reinforcement learning to safety-critical applications, we ought to provide safety guarantees during both policy training and deployment. In this work, we present theoretical results that place a bound on the probability of violating a safety property for a new task-specific policy in a model-free, episodic setting. This bound, based on a maximum policy ratio computed with respect to a 'safe' base policy, can also be applied to temporally-extended properties (beyond safety) and to robust control problems. To utilize these results, we introduce SPoRt, which provides a data-driven method for computing this bound for the base policy using the scenario approach, and includes Projected PPO, a new projection-based approach for training the task-specific policy while maintaining a user-specified bound on property violation. SPoRt thus enables users to trade off safety guarantees against task-specific performance. Complementing our theoretical results, we present experimental results demonstrating this trade-off and comparing the theoretical bound to posterior bounds derived from empirical violation rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。