让自动驾驶在无信号交叉口更安全,通过量化决策不确定性动态调整行为。
Uncertainty-Aware Safety-Critical Decision and Control for Autonomous Vehicles at Unsignalized Intersections
- 构建风险感知的集成分布强化学习,量化策略可靠性。
- 结合高阶控制屏障函数,动态降低干预频率并增强安全约束。
- 适合关注自动驾驶安全与效率平衡的研究者和开发者。
强化学习(RL)在自动驾驶决策任务中展现出潜力,但在城市道路尤其是无信号交叉口场景中仍面临重大挑战。缺乏安全约束使RL易受风险影响,且认知局限与环境随机性可能导致安全关键场景下的不可靠决策。因此,量化RL决策置信度对提升安全性至关重要。本文提出一种不确定性感知的安全关键决策与控制框架(USDC),通过构建风险感知的集成分布强化学习生成规避风险的策略,并估计不确定性以衡量策略可靠性。随后,采用高阶控制屏障函数(HOCBF)作为安全过滤器,在最小化干预策略的同时根据不确定性动态增强约束。集成评价器同时评估HOCBF与RL策略,并融入不确定性实现安全与灵活策略间的动态切换,从而平衡安全性与效率。多任务仿真测试表明,相较于基线方法,USDC在无信号交叉口场景中显著提升了安全性并维持了交通效率。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated potential in autonomous driving (AD) decision tasks. However, applying RL to urban AD, particularly in intersection scenarios, still faces significant challenges. The lack of safety constraints makes RL vulnerable to risks. Additionally, cognitive limitations and environmental randomness can lead to unreliable decisions in safety-critical scenarios. Therefore, it is essential to quantify confidence in RL decisions to improve safety. This paper proposes an Uncertainty-aware Safety-Critical Decision and Control (USDC) framework, which generates a risk-averse policy by constructing a risk-aware ensemble distributional RL, while estimating uncertainty to quantify the policy's reliability. Subsequently, a high-order control barrier function (HOCBF) is employed as a safety filter to minimize intervention policy while dynamically enhancing constraints based on uncertainty. The ensemble critics evaluate both HOCBF and RL policies, embedding uncertainty to achieve dynamic switching between safe and flexible strategies, thereby balancing safety and efficiency. Simulation tests on unsignalized intersections in multiple tasks indicate that USDC can improve safety while maintaining traffic efficiency compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。