让自动驾驶更懂伦理,显著降低对行人的事故风险。
Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
- 用伦理成本信号指导决策,融合碰撞概率与伤害严重性。
- 在两个基准上冲突频率降低25%-45%,舒适度保持稳定。
- 适合研究自动驾驶安全与伦理的学者及工程师。
自动驾驶有望大幅减少交通事故并提升交通效率,但其普及依赖于在常规和紧急操作中嵌入可信且透明的伦理推理,尤其要保护行人、骑行者等弱势道路使用者(VRUs)。本文提出一种分层安全强化学习(Safe RL)框架,在标准驾驶目标基础上引入伦理感知的成本信号。决策层采用复合伦理风险成本(结合碰撞概率与伤害严重性)训练安全强化学习代理,生成高层运动目标;动态风险敏感的优先经验回放机制增强对罕见但高危事件的学习。执行层通过多项式路径规划与PID、Stanley控制器将目标转化为平滑可行轨迹,确保精度与舒适性。我们在基于大规模真实交通数据的闭环仿真环境中训练与验证该方法,涵盖多样化车辆、骑行者与行人。结果表明,该方法在降低对他人风险的同时,维持了自身性能与舒适度。在两个交互式基准和五个随机种子下,冲突频率相比基线降低25%-45%,舒适度指标波动不超过5%。本工作为人类混合交通场景中的伦理显式安全强化学习提供了可复现基准,展示了形式化控制理论与数据驱动学习结合,推动对最脆弱道路使用者具有明确保护能力的伦理责任自治系统的发展。
原文摘要 · Abstract (English)
Autonomous vehicles hold great promise for reducing traffic fatalities and improving transportation efficiency, yet their widespread adoption hinges on embedding credible and transparent ethical reasoning into routine and emergency maneuvers, particularly to protect vulnerable road users (VRUs) such as pedestrians and cyclists. Here, we present a hierarchical Safe Reinforcement Learning (Safe RL) framework that augments standard driving objectives with ethics-aware cost signals. At the decision level, a Safe RL agent is trained using a composite ethical risk cost, combining collision probability and harm severity, to generate high-level motion targets. A dynamic, risk-sensitive Prioritized Experience Replay mechanism amplifies learning from rare but critical, high-risk events. At the execution level, polynomial path planning coupled with Proportional-Integral-Derivative (PID) and Stanley controllers translates these targets into smooth, feasible trajectories, ensuring both accuracy and comfort. We train and validate our approach on closed-loop simulation environments derived from large-scale, real-world traffic datasets encompassing diverse vehicles, cyclists, and pedestrians, and demonstrate that it outperforms baseline methods in reducing risk to others while maintaining ego performance and comfort. This work provides a reproducible benchmark for Safe RL with explicitly ethics-aware objectives in human-mixed traffic scenarios. Our results highlight the potential of combining formal control theory and data-driven learning to advance ethically accountable autonomy that explicitly protects those most at risk in urban traffic environments. Across two interactive benchmarks and five random seeds, our policy decreases conflict frequency by 25-45% compared to matched task successes while maintaining comfort metrics within 5%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。