用单智能体强化学习优化大范围交通信号,显著减少排队长度和平均通行时间。
Large-scale Regional Traffic Signal Control Based on Single-Agent Reinforcement Learning
- 单智能体架构统一控制区域信号灯,通过状态-动作-奖励设计实现全局协调。
- 在SUMO仿真中,排队长度明显下降,平均通行时间显著降低。
- 适合城市交通管理、智能信号控制研究者参考,尤其关注大规模协同优化。
在全球城市化与机动化背景下,交通拥堵已成为严重影响生活质量、环境与经济的突出问题。本文提出一种基于单智能体强化学习(RL)的区域交通信号控制(TSC)模型。与多智能体系统不同,该模型可统筹调控大范围区域内的交通信号,旨在缓解区域拥堵并最小化总通行时间。通过定义具体的状态空间(各路段队列长度与交叉口信号相位)、动作空间(先选交叉口再调整相位配比)及双目标奖励函数(缓解拥堵与最小化通行时间),构建了完整的TSC环境。实验基于SUMO交通仿真软件进行,对比无信号调整的基准情况。结果表明,该模型能有效控制拥堵:测试场景中队列长度显著减少;当奖励兼顾拥堵缓解与通行时间最小时,平均通行时间大幅下降,验证了其改善交通状况的有效性。本研究为大规模区域交通信号控制提供了新思路,对未来城市交通管理具有重要参考价值。
原文摘要 · Abstract (English)
In the context of global urbanization and motorization, traffic congestion has become a significant issue, severely affecting the quality of life, environment, and economy. This paper puts forward a single-agent reinforcement learning (RL)-based regional traffic signal control (TSC) model. Different from multi - agent systems, this model can coordinate traffic signals across a large area, with the goals of alleviating regional traffic congestion and minimizing the total travel time. The TSC environment is precisely defined through specific state space, action space, and reward functions. The state space consists of the current congestion state, which is represented by the queue lengths of each link, and the current signal phase scheme of intersections. The action space is designed to select an intersection first and then adjust its phase split. Two reward functions are meticulously crafted. One focuses on alleviating congestion and the other aims to minimize the total travel time while considering the congestion level. The experiments are carried out with the SUMO traffic simulation software. The performance of the TSC model is evaluated by comparing it with a base case where no signal-timing adjustments are made. The results show that the model can effectively control congestion. For example, the queuing length is significantly reduced in the scenarios tested. Moreover, when the reward is set to both alleviate congestion and minimize the total travel time, the average travel time is remarkably decreased, which indicates that the model can effectively improve traffic conditions. This research provides a new approach for large-scale regional traffic signal control and offers valuable insights for future urban traffic management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。