用强化学习优化红绿灯,实测延迟降低11%-32%,且训练更高效。
Evaluating the Robustness of Reinforcement Learning based Adaptive Traffic Signal Control
- 设计可支持8相位的强化学习信号控制算法,贴近实际交通设备。
- 在不同交通流量下,相比传统控制平均减少11%-32%等待时间。
- 多场景训练模型泛化能力强,适合复杂城市道路部署。
强化学习因无需环境模型、可直接从交互中学习控制策略,近年来受到自适应交通信号控制领域的关注。然而,现有研究仍面临诸多挑战:多数依赖简化信号配时结构,对模型在不同交通需求下的鲁棒性评估不足,且在微观仿真环境中训练效率偏低。本研究提出一种支持完整八相位环-障结构的强化学习信号控制算法,该算法在多种交通流量与起讫点(O-D)需求模式下进行训练与评估,并与现行主流感应式信号控制(ASC)进行对比。为评估鲁棒性,实验覆盖不同交通量及具有不同程度结构相似性的O-D模式。为提升训练效率,采用分布式异步训练架构,实现多计算节点并行仿真。案例研究显示,所提方法在各通行方向上平均延迟降低11%-32%,显著优于优化后的ASC;单一O-D模式训练的模型可在相似未见需求下良好泛化,但在差异较大场景下性能下降;而基于多样化O-D模式训练的模型则表现出强鲁棒性,在高度不同的未见场景中仍持续优于ASC。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has attracted increasing interest for adaptive traffic signal control due to its model-free ability to learn control policies directly from interaction with the traffic environment. However, several challenges remain before RL-based signal control can be considered ready for field deployment. Many existing studies rely on simplified signal timing structures, robustness of trained models under varying traffic demand conditions remains insufficiently evaluated, and runtime efficiency continues to pose challenges when training RL algorithms in traffic microscopic simulation environments. This study formulates an RL-based signal control algorithm capable of representing a full eight-phase ring-barrier configuration consistent with field signal controllers. The algorithm is trained and evaluated under varying traffic demand conditions and benchmarked against state-of-the-practice actuated signal control (ASC). To assess robustness, experiments are conducted across multiple traffic volumes and origin-destination (O-D) demand patterns with varying levels of structural similarity. To improve training efficiency, a distributed asynchronous training architecture is implemented that enables parallel simulation across multiple computing nodes. Results from a case study intersection show that the proposed RL-based signal control significantly outperforms optimized ASC, reducing average delay by 11-32% across movements. A model trained on a single O-D pattern generalizes well to similar unseen demand patterns but degrades under substantially different demand conditions. In contrast, a model trained on diverse O-D patterns demonstrates strong robustness, consistently outperforming ASC even under highly dissimilar unseen demand scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。