arXiv:2509.21745eess.SYcs.LG2025-09被引 1

用强化学习优化红绿灯,让车队排队减少近三成。

Reinforcement Learning Based Traffic Signal Design to Minimize Queue Lengths

  • 用PPO算法结合多种状态表示,自适应控制信号灯。
  • 相比传统方法,平均队列长度降低约29%。
  • 适合城市交通管理、智能信号系统研究者参考。

高效的交通信号控制(TSC)对缓解拥堵、减少出行延误、降低污染和保障道路安全至关重要。传统方法如固定周期与感应控制难以应对动态交通变化。本文提出一种基于强化学习(RL)的自适应TSC框架,采用近端策略优化(PPO)算法,以最小化所有信号相位的总队列长度为目标。为有效表征高度随机的交通状态,引入扩展状态空间、自编码器表示及受K-Planes启发的表示方法。实验基于城市交通模拟器SUMO进行,结果表明该方法在减少队列长度方面显著优于传统方法及其他主流RL方法。最优配置相较传统Webster方法实现约29%的平均队列长度降幅。此外,不同奖励函数的对比验证了基于队列的奖励设计的有效性,展示了其在可扩展、自适应城市交通管理中的潜力。

原文摘要 · Abstract (English)

Efficient traffic signal control (TSC) is crucial for reducing congestion, travel delays, pollution, and for ensuring road safety. Traditional approaches, such as fixed signal control and actuated control, often struggle to handle dynamic traffic patterns. In this study, we propose a novel adaptive TSC framework that leverages Reinforcement Learning (RL), using the Proximal Policy Optimization (PPO) algorithm, to minimize total queue lengths across all signal phases. The challenge of efficiently representing highly stochastic traffic conditions for an RL controller is addressed through multiple state representations, including an expanded state space, an autoencoder representation, and a K-Planes-inspired representation. The proposed algorithm has been implemented using the Simulation of Urban Mobility (SUMO) traffic simulator and demonstrates superior performance over both traditional methods and other conventional RL-based approaches in reducing queue lengths. The best performing configuration achieves an approximately 29% reduction in average queue lengths compared to the traditional Webster method. Furthermore, comparative evaluation of alternative reward formulations demonstrates the effectiveness of the proposed queue-based approach, showcasing the potential for scalable and adaptive urban traffic management.

交通信号控制强化学习队列优化SUMO仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。