用安全约束值生成奖励,让自动驾驶车辆在交叉路口更安全地协同决策。
Beyond Safety Filtering: Control Barrier Function-Informed Reinforcement Learning for Connected and Automated Vehicles

- 将控制屏障函数的约束值转化为强化学习奖励信号,显式引导安全行为。
- 在四车道交叉口实验中任务表现最优,且对超参数不敏感。
- 适合需要高安全性保障的多车协同自动驾驶场景研究者。
强化学习通过奖励指导学习,但奖励设计通常依赖人工启发式规则,难以调优。本文提出一种面向多智能体强化学习(MARL)的控制屏障函数(CBF)启发式奖励设计方法,将联合MARL动作下的CBF约束值转化为奖励信号,显式引导安全学习。在配备联网自动驾驶车辆的四向多车道交叉口场景中,与两种启发式奖励基线对比,结果表明该方法任务性能最高,且对奖励超参数不敏感,在测试的超参数范围内均保持稳定优异表现。实验代码及视频演示可于 https://github.com/bassamlab/SigmaRL 获取。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) uses rewards to guide learning, yet reward design is typically hand-crafted using heuristics that can be difficult to tune. We propose a Control Barrier Function (CBF)-informed reward design for Multi-Agent RL (MARL) that converts CBF constraint values under joint MARL actions into a reward signal that explicitly guides safe learning. We compare against two heuristic reward baselines in a four-way multi-lane intersection with connected and automated vehicles. Results show that our method achieves the highest task performance and is less sensitive to reward hyperparameters, yielding consistently strong performance across the tested hyperparameter range. Code for reproducing the experimental results and a video demonstration are available at https://github.com/bassamlab/SigmaRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。