arXiv:2509.07412cs.RO2025-09

改进PPO强化学习,让自动驾驶更安全高效

Attention and Risk-Aware Decision Framework for Safe Autonomous Driving

  • 引入风险感知机制与注意力网络,聚焦高危区域
  • 训练效率提升,碰撞率降低,峰值奖励更高
  • 适合自动驾驶安全控制研究者与工程师

自动驾驶因具备全无人驾驶潜力而备受关注。模型驱动与学习驱动方法被广泛采用,但前者难以应对突发状况,后者如近端策略优化(PPO)存在训练效果差、长序列训练效率低的问题,且不良训练结果等同于驾驶碰撞。为此,本文提出改进PPO算法,引入风险感知机制、风险注意力决策网络、平衡奖励函数与安全辅助机制。风险感知机制突出潜在碰撞区域,促进安全驾驶学习;平衡奖励函数根据周围车辆数动态调整奖励,提升策略探索效率;风险注意力网络对输入图像的高风险区域实施通道与空间注意力;安全辅助机制在变道与车道保持过程中监督并阻止高风险动作。在物理引擎上的仿真结果表明,该算法在多种交通流场景下优于基准方法,实现更高峰值奖励、更短训练时间,并减少在高风险区域的停留时间。

原文摘要 · Abstract (English)

Autonomous driving has attracted great interest due to its potential capability in full-unsupervised driving. Model-based and learning-based methods are widely used in autonomous driving. Model-based methods rely on pre-defined models of the environment and may struggle with unforeseen events. Proximal policy optimization (PPO), an advanced learning-based method, can adapt to the above limits by learning from interactions with the environment. However, existing PPO faces challenges with poor training results, and low training efficiency in long sequences. Moreover, the poor training results are equivalent to collisions in driving tasks. To solve these issues, this paper develops an improved PPO by introducing the risk-aware mechanism, a risk-attention decision network, a balanced reward function, and a safety-assisted mechanism. The risk-aware mechanism focuses on highlighting areas with potential collisions, facilitating safe-driving learning of the PPO. The balanced reward function adjusts rewards based on the number of surrounding vehicles, promoting efficient exploration of the control strategy during training. Additionally, the risk-attention network enhances the PPO to hold channel and spatial attention for the high-risk areas of input images. Moreover, the safety-assisted mechanism supervises and prevents the actions with risks of collisions during the lane keeping and lane changing. Simulation results on a physical engine demonstrate that the proposed algorithm outperforms benchmark algorithms in collision avoidance, achieving higher peak reward with less training time, and shorter driving time remaining on the risky areas among multiple testing traffic flow scenarios.

自动驾驶强化学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。