用强化学习动态调信号,抗需求波动,降排队长度。
Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations
- 单智能体集中决策,用邻接矩阵融合路网、排队和信号状态。
- 在10%~30%需求波动下,队列长度显著降低,抗干扰能力强。
- 适合城市交通管控、智慧交管系统研发人员参考。
交通拥堵主要由交叉口排队引发,严重影响城市生活质量、安全、环境与经济效率。传统交通信号控制(TSC)模型难以捕捉真实交通的复杂性与动态性。本文提出一种新型单智能体强化学习框架,用于区域自适应交通信号控制,通过中心化决策避免多智能体协作复杂性。模型采用邻接矩阵统一编码道路网络拓扑、基于探针车数据的实时队列状态及当前信号配时参数。借助DreamerV3世界模型高效学习能力,智能体学习控制策略:动作依次选择交叉口并调整其信号相位分配,实现对车流进出的调节,类比反馈控制系统。奖励设计聚焦队列消散,直接将队列长度等拥堵指标与控制动作关联。在SUMO仿真中,面对10%、20%、30%的起讫点(OD)需求波动,该框架展现出强鲁棒性,显著降低队列长度。本研究为兼容探针车技术的智能交通控制提供新范式。未来工作将引入训练阶段的随机OD需求波动,并探索应对突发事件的区域优化机制。
原文摘要 · Abstract (English)
Traffic congestion, primarily driven by intersection queuing, significantly impacts urban living standards, safety, environmental quality, and economic efficiency. While Traffic Signal Control (TSC) systems hold potential for congestion mitigation, traditional optimization models often fail to capture real-world traffic complexity and dynamics. This study introduces a novel single-agent reinforcement learning (RL) framework for regional adaptive TSC, circumventing the coordination complexities inherent in multi-agent systems through a centralized decision-making paradigm. The model employs an adjacency matrix to unify the encoding of road network topology, real-time queue states derived from probe vehicle data, and current signal timing parameters. Leveraging the efficient learning capabilities of the DreamerV3 world model, the agent learns control policies where actions sequentially select intersections and adjust their signal phase splits to regulate traffic inflow/outflow, analogous to a feedback control system. Reward design prioritizes queue dissipation, directly linking congestion metrics (queue length) to control actions. Simulation experiments conducted in SUMO demonstrate the model's effectiveness: under inference scenarios with multi-level (10%, 20%, 30%) Origin-Destination (OD) demand fluctuations, the framework exhibits robust anti-fluctuation capability and significantly reduces queue lengths. This work establishes a new paradigm for intelligent traffic control compatible with probe vehicle technology. Future research will focus on enhancing practical applicability by incorporating stochastic OD demand fluctuations during training and exploring regional optimization mechanisms for contingency events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。