通过预测环境变化提前设置安全约束,避免自动驾驶在动态场景中出险。
Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning

- 基于观测推断环境状态并预测其演化,生成前瞻性安全约束
- 在高速公路、交叉口等场景中碰撞率显著降低,跨场景泛化性强
- 适合动态环境下的自动驾驶等对安全性要求高的强化学习任务
在非平稳环境中保障强化学习的安全性,需在危险行为发生前预判风险变化。现有方法多依赖设计时设定或执行中被动更新的安全约束,假设约束长期有效,但在环境上下文与道路布局持续演变的场景下可能失效。本文提出一种基于上下文预测的主动安全约束生成框架:从观测中推断隐含环境状态,预测其未来演化,并构建适配预期条件的安全约束。该方法使智能体能主动规避潜在危险区域,而非事后响应。我们在具有结构化上下文变化的驾驶环境中进行评估,涵盖不同非平稳强度及未见的道路布局(包括高速公路、交叉口、赛道)。结果表明,主动约束生成在已知与未知的非平稳条件下均显著减少碰撞,且在各类新布局中保持良好性能,同时维持可接受的任务表现。这表明基于上下文的约束生成是应对非平稳强化学习安全挑战的可行方案。
原文摘要 · Abstract (English)
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before they lead to unsafe behavior. Existing approaches typically rely on safety constraints defined at design time or updated reactively during execution, assuming that such constraints remain valid over time. However, in nonstationary environments with evolving contexts and changing driving layouts, these assumptions may fail. We propose a framework for proactive safety constraint generation based on context forecasting. The approach infers latent environmental context from observations, predicts its future evolution, and constructs safety constraints adapted to anticipated conditions. This enables the agent to proactively avoid unsafe regions instead of reacting only after safety violations occur. We evaluate the method in driving environments with structured context variation. The experiments include a sweep over nonstationarity intensities and additional held-out driving layouts, including highway, intersection, and racetrack scenarios. Results show that proactive constraint generation substantially reduces collisions under both seen and out-of-training nonstationarity intensities and generally remains effective across held-out driving layouts while maintaining usable task performance. These findings suggest that context-based constraint generation is a promising approach for safe reinforcement learning under nonstationarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。