提出动态残差安全强化学习框架,提升多智能体安全决策的实时性与效率。
Dynamic Residual Safe Reinforcement Learning for Multi-Agent Safety-Critical Scenarios Decision-Making
- 基于弱到强安全修正机制,动态调整安全边界,无需人工规则。
- 碰撞率降低92.17%,安全模型仅占主模型27%参数量。
- 适合自动驾驶等高安全要求场景,兼顾安全性与策略灵活性。
在多智能体安全关键场景中,传统自动驾驶框架难以兼顾安全约束与任务性能,难以实时量化动态交互风险,依赖大量人工规则,导致计算效率低且策略保守。为此,我们提出一种基于增强型网络马尔可夫决策过程的动态残差安全强化学习(DRS-RL)框架。首次将弱到强理论引入多智能体决策,通过弱到强安全修正范式实现轻量级动态安全边界校准。基于多智能体动态冲突区模型,准确捕捉异质交通参与者间的时空耦合风险,突破传统几何规则的静态约束。此外,引入风险感知优先经验回放机制,通过将风险映射为采样概率,缓解数据分布偏差。实验表明,该方法在安全性、效率与舒适性上显著优于传统强化学习算法:碰撞率最高降低92.17%,且安全模型仅占主模型27%参数。该框架适用于复杂交通环境下的自主决策。
原文摘要 · Abstract (English)
In multi-agent safety-critical scenarios, traditional autonomous driving frameworks face significant challenges in balancing safety constraints and task performance. These frameworks struggle to quantify dynamic interaction risks in real-time and depend heavily on manual rules, resulting in low computational efficiency and conservative strategies. To address these limitations, we propose a Dynamic Residual Safe Reinforcement Learning (DRS-RL) framework grounded in a safety-enhanced networked Markov decision process. It's the first time that the weak-to-strong theory is introduced into multi-agent decision-making, enabling lightweight dynamic calibration of safety boundaries via a weak-to-strong safety correction paradigm. Based on the multi-agent dynamic conflict zone model, our framework accurately captures spatiotemporal coupling risks among heterogeneous traffic participants and surpasses the static constraints of conventional geometric rules. Moreover, a risk-aware prioritized experience replay mechanism mitigates data distribution bias by mapping risk to sampling probability. Experimental results reveal that the proposed method significantly outperforms traditional RL algorithms in safety, efficiency, and comfort. Specifically, it reduces the collision rate by up to 92.17%, while the safety model accounts for merely 27% of the main model's parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。