用预测安全风险的模型,让自动驾驶更安全高效。
Learning Safe Autonomous Driving Policies Using Predictive Safety Representations
- 引入可预测未来违规行为的安全表征,指导驾驶策略优化
- 在真实数据集上实现成功率提升(效应量0.65-0.86)与成本降低
- 增强对观测噪声的鲁棒性,跨数据集泛化能力更强
安全强化学习(SafeRL)是自动驾驶中的关键范式,要求在严格安全约束下优化性能。此双重目标导致根本性矛盾:过度保守降低效率,激进探索则威胁安全。本文研究了基于预测安全表征的安全策略学习(SRPL)框架在真实场景中的适用性。在Waymo Open Motion Dataset(WOMD)和NuPlan数据集上的系统实验表明,SRPL可显著改善奖励-安全权衡,成功率提升(效应量r=0.65–0.86)且成本下降(效应量r=0.70–0.83),p<0.05。其效果依赖于底层策略优化器与数据分布。结果还显示,预测性安全表征显著提升对观测噪声的鲁棒性。零样本跨数据集评估中,SRPL增强的智能体表现出优于非SRPL方法的泛化能力。这些发现共同证明,预测性安全表征能有效强化自动驾驶中的安全强化学习。
原文摘要 · Abstract (English)
Safe reinforcement learning (SafeRL) is a prominent paradigm for autonomous driving, where agents are required to optimize performance under strict safety requirements. This dual objective creates a fundamental tension, as overly conservative policies limit driving efficiency while aggressive exploration risks safety violations. The Safety Representations for Safer Policy Learning (SRPL) framework addresses this challenge by equipping agents with a predictive model of future constraint violations and has shown promise in controlled environments. This paper investigates whether SRPL extends to real-world autonomous driving scenarios. Systematic experiments on the Waymo Open Motion Dataset (WOMD) and NuPlan demonstrate that SRPL can improve the reward-safety tradeoff, achieving statistically significant improvements in success rate (effect sizes r = 0.65-0.86) and cost reduction (effect sizes r = 0.70-0.83), with p < 0.05 for observed improvements. However, its effectiveness depends on the underlying policy optimizer and the dataset distribution. The results further show that predictive safety representations play a critical role in improving robustness to observation noise. Additionally, in zero-shot cross-dataset evaluation, SRPL-augmented agents demonstrate improved generalization compared to non-SRPL methods. These findings collectively demonstrate the potential of predictive safety representations to strengthen SafeRL for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。