arXiv:2605.18842cs.LG2026-05

提出自适应安全约束框架,让智能体在环境变化时持续安全决策。

Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints

论文配图:Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints
图 1 · 摘自论文原文
  • 根据环境上下文动态调整安全阈值
  • 环境变化过快时自动收紧安全限制,减少违规
  • 适合需要长期安全运行的自动驾驶等场景

非平稳环境中安全强化学习需要能随环境变化自适应的安全机制。传统方法多假设约束固定或环境稳定,但在分布偏移下可能失效。本文提出LILAC+框架,融合三种自适应安全机制:基于上下文的安全约束、适应速度约束和预算到状态的安全执行。上下文约束利用推断和预测的环境信息动态调整安全要求;适应速度约束在环境变化速率超过智能体安全适应能力时收紧安全限制;预算到状态机制将累积安全要求转化为可实时执行的局部状态约束。在模拟驾驶场景中评估显示,该框架在平稳、已知非平稳和未知非平稳条件下均显著降低安全违规,同时保持与无约束及固定约束基线相当的任务性能。结果表明,安全持续强化学习需依赖响应当前状态、预测环境上下文、适应需求和剩余安全预算的自适应约束机制。

原文摘要 · Abstract (English)

Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinforcement learning methods often assume fixed constraints or stable environmental conditions, which can become inadequate under distribution shift. We propose LILAC+, a framework for safe continual reinforcement learning under nonstationarity that combines three adaptive safety mechanisms: context-based safety constraints, adaptation-speed constraints, and budget-to-state safety enforcement. Context-based constraints adjust safety requirements using inferred and predicted environmental context. Adaptation-speed constraints tighten safety requirements when the rate of environmental change exceeds the agent's ability to adapt safely. Budget-to-state enforcement converts cumulative safety requirements into local state-level control constraints that can be enforced at decision time. Together, these mechanisms provide a unified approach for proactive and reactive safety adaptation in continual reinforcement learning. We evaluate the framework in simulated driving environments under stationary, seen nonstationary, and unseen nonstationary conditions. The results show that adaptive safety constraints substantially reduce safety violations under distribution shift while maintaining competitive task performance compared with unconstrained and fixed-constraint baselines. These findings suggest that safe continual reinforcement learning requires adaptive constraint mechanisms that respond not only to current state information but also to predicted environmental context, adaptation demand, and remaining safety budget.

强化学习安全控制非平稳性自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。