混合模型与强化学习协同控制信号灯,自适应应对不同交通需求。
A Hierarchical Signal Coordination and Control System Using a Hybrid Model-based and Reinforcement Learning Approach
- 分层设计:高层选策略,中层定相位约束,底层用强化学习调信号。
- 重载时混合MFC吞吐量最高,轻载时混合GWC减少停驶且保持车流顺畅。
- 适合城市道路网信号优化,尤其对动态交通有强适应性。
城市干道信号控制需兼顾主干道车流连贯性与局部路口需求变化。本文提出一种分层信号协调控制方案,融合模型优化与强化学习。系统包含:(i) 高层协调器(HLC)根据观测和预测需求选择协调策略;(ii) 干道协调器从所选策略(最大通行量协调MFC或绿波协调GWC)推导相位约束;(iii) 混合信号代理(HSAs)通过带动作掩码的强化学习确定信号相位。采用近端策略优化(PPO)进行分层强化学习训练。底层训练三种策略:面向MFC、面向GWC及纯代理控制(PAC)。高层HLC则通过多目标奖励平衡路网级与全局性能,实现动态策略切换。在SUMO-RLlib平台评估显示:混合MFC在高需求下最大化通行量;混合GWC在各类条件下持续降低干道停驶次数并维持车流连贯,但可能降低整体效率;PAC在中等需求下改善路网平均出行时间,但在高需求下表现较弱。分层结构支持自适应策略选择,在所有需求水平下均表现出鲁棒性能。
原文摘要 · Abstract (English)
Signal control in urban corridors faces the dual challenge of maintaining arterial traffic progression while adapting to demand variations at local intersections. We propose a hierarchical traffic signal coordination and control scheme that integrates model-based optimization with reinforcement learning. The system consists of: (i) a High-Level Coordinator (HLC) that selects coordination strategies based on observed and predicted demand; (ii) a Corridor Coordinator that derives phase constraints from the selected strategy-either Max-Flow Coordination (MFC) or Green-Wave Coordination (GWC); and (iii) Hybrid Signal Agents (HSAs) that determine signal phases via reinforcement learning with action masking to enforce feasibility. Hierarchical reinforcement learning with Proximal Policy Optimization (PPO) is used to train HSA and HLC policies. At the lower level, three HSA policies-MFC-aware, GWC-aware, and pure agent control (PAC) are trained in conjunction with their respective coordination strategies. At the higher level, the HLC is trained to dynamically switch strategies using a multi-objective reward balancing corridor-level and network-wide performance. The proposed scheme was developed and evaluated on a SUMO-RLlib platform. Case results show that hybrid MFC maximizes throughput under heavy demand; hybrid GWC consistently minimizes arterial stops and maintains progression across diverse traffic conditions but can reduce network-wide efficiency; and PAC improves network-wide travel time in moderate demand but is less effective under heavy demand. The hierarchical design enables adaptive strategy selection, achieving robust performance across all demand levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。