arXiv:2505.08896cs.AIcs.RO2025-05被引 1

用深度强化学习优化红绿灯路口车辆纵向控制,提升效率与舒适性。

Deep reinforcement learning-based longitudinal control strategy for automated vehicles at signalised intersections

  • 基于DRL设计奖励函数,融合跟车效率、黄灯决策与加减速不对称性。
  • 模型在真实轨迹上测试,头间距更小、急加速减速度更低。
  • 适用于智能驾驶车辆在复杂路口场景下的安全高效控制。

针对信号灯交叉口自主车辆控制的复杂决策挑战,本文提出一种基于深度强化学习(DRL)的纵向控制策略。设计了综合奖励函数,重点包含:(i) 基于车距的效率奖励,(ii) 黄灯阶段的决策准则,(iii) 加/减速的非对称响应,同时保留传统安全与舒适性要求。该奖励函数结合Deep Deterministic Policy Gradient (DDPG) 和 Soft Actor-Critic (SAC) 两种主流DRL算法,可处理加速度/减速度的连续动作空间。模型在真实前车轨迹与基于奥恩斯坦-乌伦贝克(Ornstein-Uhlenbeck, OU)过程生成的模拟轨迹组合数据上训练。通过累积分布函数(CDF)图对比真实轨迹,结果表明:相比人类驾驶车辆,该模型在不牺牲安全性的前提下,实现了更小的车距头间距(即更高效率)和更低的急动度(jerk)。进一步评估在多种高风险场景下的鲁棒性,包括跟车与信号灯合规性,两种模型均成功应对,其中DDPG模型表现出更平滑的动作轨迹。总体验证了基于DRL的纵向控制策略可有效提升交通安全性、效率与乘坐舒适性。

原文摘要 · Abstract (English)

Developing an autonomous vehicle control strategy for signalised intersections (SI) is one of the challenging tasks due to its inherently complex decision-making process. This study proposes a Deep Reinforcement Learning (DRL) based longitudinal vehicle control strategy at SI. A comprehensive reward function has been formulated with a particular focus on (i) distance headway-based efficiency reward, (ii) decision-making criteria during amber light, and (iii) asymmetric acceleration/ deceleration response, along with the traditional safety and comfort criteria. This reward function has been incorporated with two popular DRL algorithms, Deep Deterministic Policy Gradient (DDPG) and Soft-Actor Critic (SAC), which can handle the continuous action space of acceleration/deceleration. The proposed models have been trained on the combination of real-world leader vehicle (LV) trajectories and simulated trajectories generated using the Ornstein-Uhlenbeck (OU) process. The overall performance of the proposed models has been tested using Cumulative Distribution Function (CDF) plots and compared with the real-world trajectory data. The results show that the RL models successfully maintain lower distance headway (i.e., higher efficiency) and jerk compared to human-driven vehicles without compromising safety. Further, to assess the robustness of the proposed models, we evaluated the model performance on diverse safety-critical scenarios, in terms of car-following and traffic signal compliance. Both DDPG and SAC models successfully handled the critical scenarios, while the DDPG model showed smoother action profiles compared to the SAC model. Overall, the results confirm that DRL-based longitudinal vehicle control strategy at SI can help to improve traffic safety, efficiency, and comfort.

自动驾驶强化学习交通控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。