arXiv:2604.16702cs.RO2026-04

用赛车参数化强化学习实现高效自动驾驶避撞,性能远超传统方法。

Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning

论文配图:Autonomous Vehicle Collision Avoidance With Racing Parameterized Deep Reinforcement Learning
图 1 · 摘自论文原文
  • 基于赛车场景的参数化深度强化学习,无需几何参考轨迹。
  • 在三种交叉口场景中避撞成功率比传统方法高10%,计算量减少64倍。
  • 适合需要低延迟、高动态响应的自动驾驶系统研发人员。

道路交通事故是全球主要致死原因。美国94%的事故由人为错误引发,每年导致逾7000名行人死亡,经济损失达5000亿美元。具备紧急避撞功能的自动驾驶车辆(AV)在极端天气和网络攻击下,以高频运行于车辆动力学极限,同时满足非线性运动精度与计算效率双重约束,可显著提升安全性。本文提出一种基于赛车超车场景的参数化深度强化学习(DRL)避撞策略,在无显式几何参考轨迹引导下,通过物理感知、仿真器感知的奖励函数编码非线性车辆动力学。评估了默认单向与反向行驶两种策略,两者在三个交叉路口避撞场景中均显著优于模型预测控制与人工势场法(MPC-APF)基线,且实现零样本迁移至比例缩放的硬件平台。计算量仅需基线31倍浮点运算,推理延迟降低64倍。反向行驶策略在正面碰撞中表现优于默认策略30%,较基线提升50%;侧面碰撞中与默认策略持平,两者避撞成功率均比数值最优控制高10%。

原文摘要 · Abstract (English)

Road traffic accidents are a leading cause of fatalities worldwide. In the US, human error causes 94% of crashes, resulting in excess of 7,000 pedestrian fatalities and $500 billion in costs annually. Autonomous Vehicles (AVs) with emergency collision avoidance systems that operate at the limits of vehicle dynamics at a high frequency, a dual constraint of nonlinear kinodynamic accuracy and computational efficiency, further enhance safety benefits during adverse weather and cybersecurity breaches, and to evade dangerous human driving when AVs and human drivers share roads. This paper parameterizes a Deep Reinforcement Learning (DRL) collision avoidance policy Out-Of-Distribution (OOD) utilizing race car overtaking, without explicit geometric mimicry reference trajectory guidance, in simulation, with a physics-informed, simulator exploit-aware reward to encode nonlinear vehicle kinodynamics. Two policies are evaluated, a default uni-direction and a reversed heading variant that navigates in the opposite direction to other cars, which both consistently outperform a Model Predictive Control and Artificial Potential Function (MPC-APF) baseline, with zero-shot transfer to proportionally scaled hardware, across three intersection collision scenarios, at 31x fewer Floating Point Operations (FLOPS) and 64x lower inference latency. The reversed heading policy outperforms the default racing overtaking policy in head-to-head collisions by 30% and the baseline by 50%, and matches the former in side collisions, where both DRL policies evade 10% greater than numerical optimal control.

自动驾驶强化学习避撞系统实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。