提出安全持续强化学习方法,解决非线性系统在变化条件下的安全与记忆难题。
On the Design of Safe Continual RL Methods for Control of Nonlinear Systems
- 设计奖励重塑策略,让模型优先保留安全与任务性能。
- 实验表明现有方法在速度约束和关节故障下会违反安全约束。
- 适用于无人机、机器人等需要长期安全运行的复杂控制系统。
强化学习已成功应用于无人机和机器人控制任务。近年来,安全强化学习被提出以确保在工业和关键任务系统中闭环运行的安全性。然而,当系统运行条件改变(如未知故障发生)时,典型的安全强化学习算法无法在保留已有知识的同时适应新情况。持续强化学习算法虽能应对此问题,但其对系统安全性的影响尚未充分研究。本文探讨安全与持续强化学习的交叉领域。首先,我们通过实验证明,一种流行的持续强化学习算法——在线弹性权重巩固,在受速度约束且存在突然关节失效非平稳性的MuJoCo HalfCheetah和Ant环境中,无法满足安全约束。其次,我们发现使用约束策略优化训练的智能体在持续学习场景中会出现灾难性遗忘。为此,我们提出一种简单的奖励重塑方法,使弹性权重巩固在非线性、非平稳动态系统中同时优先记住安全性和任务性能。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms have been successfully applied to control tasks associated with unmanned aerial vehicles and robotics. In recent years, safe RL has been proposed to allow the safe execution of RL algorithms in industrial and mission-critical systems that operate in closed loops. However, if the system operating conditions change, such as when an unknown fault occurs in the system, typical safe RL algorithms are unable to adapt while retaining past knowledge. Continual reinforcement learning algorithms have been proposed to address this issue. However, the impact of continual adaptation on the system's safety is an understudied problem. In this paper, we study the intersection of safe and continual RL. First, we empirically demonstrate that a popular continual RL algorithm, online elastic weight consolidation, is unable to satisfy safety constraints in non-linear systems subject to varying operating conditions. Specifically, we study the MuJoCo HalfCheetah and Ant environments with velocity constraints and sudden joint loss non-stationarity. Then, we show that an agent trained using constrained policy optimization, a safe RL algorithm, experiences catastrophic forgetting in continual learning settings. With this in mind, we explore a simple reward-shaping method to ensure that elastic weight consolidation prioritizes remembering both safety and task performance for safety-constrained, non-linear, and non-stationary dynamical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。