让强化学习在动态变化中既安全又持续适应,解决物理系统长期运行的难题。
Safe Continual Reinforcement Learning in Non-stationary Environments

- 设计三个安全关键的持续适应基准环境,评估不同方法表现。
- 发现现有方法难以同时兼顾安全约束与避免灾难性遗忘。
- 提出正则化策略缓解冲突,适合需要长期自主运行的控制场景。
强化学习(RL)为缺乏精确物理模型的复杂系统提供了一种数据驱动的控制器设计范式;然而,多数面向控制的RL方法假设环境静态,因此在真实非平稳环境中表现不佳,因系统动态和运行条件可能意外变化。此外,作用于物理环境的RL控制器必须在整个学习与执行阶段满足安全约束,不可接受适应过程中的瞬时违规。尽管持续强化学习和安全强化学习分别解决了非平稳性和安全性问题,但两者的交集仍研究不足。本文系统研究安全持续强化学习,引入三个捕捉安全关键持续适应的基准环境,并评估来自安全RL、持续RL及其组合的代表性方法。实验结果揭示了在非平稳动态下,维持安全约束与防止灾难性遗忘之间存在根本矛盾,现有方法普遍无法同时实现两个目标。为此,我们考察基于正则化的策略,部分缓解该权衡并刻画其优劣。最后,指出关键开放挑战与研究方向,以发展可持久自主运行于变化环境的安全鲁棒学习型控制器。
原文摘要 · Abstract (English)
Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are unavailable; however, most existing control-oriented RL methods assume stationarity and, therefore, struggle in real-world non-stationary deployments where system dynamics and operating conditions can change unexpectedly. Moreover, RL controllers acting in physical environments must satisfy safety constraints throughout their learning and execution phases, rendering transient violations during adaptation unacceptable. Although continual RL and safe RL have each addressed non-stationarity and safety, respectively, their intersection remains comparatively unexplored, motivating the study of safe continual RL algorithms that can adapt over the system's lifetime while preserving safety. In this work, we systematically investigate safe continual reinforcement learning by introducing three benchmark environments that capture safety-critical continual adaptation and by evaluating representative approaches from safe RL, continual RL, and their combinations. Our empirical results reveal a fundamental tension between maintaining safety constraints and preventing catastrophic forgetting under non-stationary dynamics, with existing methods generally failing to achieve both objectives simultaneously. To address this shortcoming, we examine regularization-based strategies that partially mitigate this trade-off and characterize their benefits and limitations. Finally, we outline key open challenges and research directions toward developing safe, resilient learning-based controllers capable of sustained autonomous operation in changing environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。