提出用调整速度约束提升非平稳强化学习的安全性
Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

- 以适应可行性定义安全,通过上下文预测判断是否能及时调整
- 实验显示该方法在环境突变时显著减少短期安全违规
- 适合需要实时安全防护的自动驾驶等动态系统
在非平稳强化学习中,确保安全的关键在于判断学习系统能否在规定恢复期内安全适应预期环境变化。现有方法多假设环境平稳,未将调整速度作为安全约束。当环境随时间演变时,延迟适应可能导致暂时不安全行为。本文提出将调整速度作为非平稳强化学习的安全约束,核心思想是:若为保持安全所需适应量超过系统的校准恢复能力,则未来状态可能变得不安全。所提框架利用学习到的上下文表示和短时上下文预测,估计适应需求,并与代理可实现的适应能力进行比较。当预测适应需求超出校准恢复能力时,框架会主动收紧允许动作集,并激活动作级防护机制,提前降低不安全行为风险。在非平稳驾驶环境中实验表明,该方法主要减少了与上下文变化对齐的短时窗口内的安全违规。消融研究进一步显示,防护机制对峰值与尾部风险抑制更保守,而优化级调整还能额外降低短时切换条件下的违规率。结果支持适应可行性作为非平稳强化学习中的实用安全原则,并证明主动干预可在环境变化期提升安全性。
原文摘要 · Abstract (English)
Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement learning methods typically assume stationary environments and do not explicitly consider adaptation speed as a safety concern. However, when environments evolve over time, delayed adaptation may result in transient unsafe behavior. This paper proposes adjustment speed as a safety constraint for nonstationary reinforcement learning. The central idea is to define safety in terms of adaptation feasibility: future states or regions may become unsafe when the adaptation required to remain safe exceeds the learning system's calibrated recovery capacity. The proposed framework uses learned context representations and short-horizon context forecasts to estimate adaptation demand and compare it with the agent's achievable adaptation capacity. When predicted adaptation demand exceeds the calibrated recovery capacity, the framework proactively tightens the admissible action set and activates an action-level shield to reduce unsafe behavior before violations occur. Experiments in a nonstationary driving environment show that the proposed approach primarily reduces safety violations in short-horizon windows aligned with context changes. Ablation studies further show that shielding is more conservative for peak- and tail-risk suppression, while optimization-level adjustment provides additional reductions in short-horizon switch-conditioned violations. These results support adaptation feasibility as a practical safety principle for reinforcement learning under nonstationarity and demonstrate that proactive intervention can improve safety during periods of environmental change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。