arXiv:2608.26860cs.LGcs.AI2026-08

用强化学习优化自动驾驶车队变道,兼顾安全与效率

Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic

论文配图:Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
图 1 · 摘自论文原文
  • 采用PPO算法并加入风险惩罚项提升决策安全性
  • 成功率达98%,碰撞率低于1%,但需更多决策步数
  • 适合研究智能交通系统或自动驾驶控制的学者

联网自动驾驶车辆(CAV)编队可显著提升道路安全与通行能力,但在混合交通环境中,因人类驾驶行为差异和不确定性,编队控制面临挑战。本文提出一个通用建模与仿真框架,用于研究CAV编队变道行为,并对比基于深度强化学习(DRL)的控制算法。在与人类车辆共存的复杂交通场景中,目标是实现安全高效的变道操作。通过将风险惩罚纳入奖励函数,或引入外部安全控制器约束策略,评估了Deep Q-Network(DQN)、Double Deep Q-Network(DDQN)和Proximal Policy Optimization(PPO)三种算法。结果表明,PPO表现最优,成功率达约98%,碰撞率低于1%,主要得益于奖励函数中的风险惩罚机制。然而,其性能提升以增加决策步数为代价,揭示了安全、效率与决策速度间的权衡。外部安全控制器虽能有效防止碰撞,但可能降低变道效率。研究强调,在设计混合交通环境下基于强化学习的编队控制时,必须协同考虑安全与效率。

原文摘要 · Abstract (English)

Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety and traffic capacity. However, platoon control in real-world traffic is challenging due to uncertainty and heterogeneous driving behaviors. Reinforcement learning (RL) has strong potential for addressing such control problems, but its practical deployment raises challenges related to safety and learning efficiency. This paper proposes a generic modeling and simulation framework for investigating CAV platoon joining maneuvers and comparing deep reinforcement learning (DRL)-based control algorithms. The problem is particularly challenging in mixed-traffic environments, where CAVs coexist with human-driven vehicles exhibiting heterogeneous longitudinal and lateral behaviors. The objective is to achieve safe and efficient joining maneuvers by either incorporating penalties for risky behaviors into the learning process or using an external safety controller to constrain the learned policy. An agent-based modeling framework coupled with the Simulation of Urban MObility (SUMO) simulator is used to evaluate Deep Q-Network (DQN), Double Deep Q-Network (DDQN), and Proximal Policy Optimization (PPO). Results show that PPO outperforms DQN and DDQN, achieving a joining success rate of approximately 98 % and a collision rate below 1 %, largely due to risk-related penalties incorporated into the reward function. However, this improved performance requires more decision steps to complete the maneuver, revealing a trade-off between safety, joining effectiveness, and decision efficiency. An external safety controller effectively prevents collisions, although its interventions may reduce joining efficiency. The results highlight the importance of jointly considering safety and efficiency when designing RL-based controllers for CAV platoon joining in mixed traffic.

自动驾驶强化学习交通控制车队编队

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。