arXiv:2509.03118cs.LGcs.AI2025-09被引 1

分层强化学习优化信号灯周期,兼顾公平与效率。

A Hierarchical Deep Reinforcement Learning Framework for Traffic Signal Control with Predictable Cycle Planning

  • 高层决策方向时长,低层细化直行左转分配。
  • 实测显示比现有方法平均减少23%车辆等待时间。
  • 适合城市交叉口信号优化,尤其关注驾驶体验的场景。

深度强化学习(DRL)在交通信号控制(TSC)中因能从复杂交通环境中学习自适应策略而广受欢迎。现有方法主要分为‘选相位’和‘切换’两类:前者虽可动态调整相位顺序,但易导致司机预期混乱,影响安全;后者虽保持相位顺序可预测,却常造成某些方向或转向过度延长,其他被忽略,引发不公平与低效。本文提出一种名为深度分层周期规划器(DHCP)的DRL模型,实现信号周期的分层分配:高层智能体根据整体交通状态决定南北(NS)与东西(EW)方向的周期时间分配;低层智能体则在每个主方向内进一步划分直行与左转的时长,提升灵活性。我们在真实与合成路网及多组真实与合成交通流上测试该模型。实验结果表明,其在所有数据集上均优于基线方法,显著降低车辆等待时间并提升通行效率。

原文摘要 · Abstract (English)

Deep reinforcement learning (DRL) has become a popular approach in traffic signal control (TSC) due to its ability to learn adaptive policies from complex traffic environments. Within DRL-based TSC methods, two primary control paradigms are ``choose phase" and ``switch" strategies. Although the agent in the choose phase paradigm selects the next active phase adaptively, this paradigm may result in unexpected phase sequences for drivers, disrupting their anticipation and potentially compromising safety at intersections. Meanwhile, the switch paradigm allows the agent to decide whether to switch to the next predefined phase or extend the current phase. While this structure maintains a more predictable order, it can lead to unfair and inefficient phase allocations, as certain movements may be extended disproportionately while others are neglected. In this paper, we propose a DRL model, named Deep Hierarchical Cycle Planner (DHCP), to allocate the traffic signal cycle duration hierarchically. A high-level agent first determines the split of the total cycle time between the North-South (NS) and East-West (EW) directions based on the overall traffic state. Then, a low-level agent further divides the allocated duration within each major direction between straight and left-turn movements, enabling more flexible durations for the two movements. We test our model on both real and synthetic road networks, along with multiple sets of real and synthetic traffic flows. Empirical results show our model achieves the best performance over all datasets against baselines.

交通信号强化学习分层控制城市交通

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。