arXiv:2602.17068cs.LGcs.SY2026-02

用双阶段超图强化学习,让交通信号更优兼顾有轨电车与公交车。

Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control

  • 构建时空双阶段超图捕捉行人与车辆交互关系
  • 电车等待时间下降显著,公交表现随场景波动
  • 适合研究智能交通系统或城市规划的工程师

在走廊网络中,以人为本的交通信号控制需越来越多考虑多模式出行者,尤其是高载客量的公共交通,而非仅关注车辆性能。本文提出基于时空双阶段超图的多智能体强化学习框架STDSH-MARL,采用集中训练、分散执行范式。该方法通过新颖的双阶段超图注意力机制,建模空间与时间超边上的交互关系,以捕捉时空依赖性。同时引入混合离散动作空间,联合决策下一信号相位及其绿灯时长,实现更灵活的信号配时。在五种交通场景下的走廊网络实验表明,STDSH-MARL实现了出色的综合多模式性能:电车等待时间显著且相对稳定地降低,而公交车等待时间改善则因场景不同而异。结果凸显了整体网络效率、电车优先权与公交服务质量之间的权衡。相较于现有先进基线方法,本方法整体表现更优。消融实验进一步验证各组件贡献,其中时间超边被确认为推动性能提升的关键因素。

原文摘要 · Abstract (English)

Human-centric traffic signal control in corridor networks must increasingly account for multimodal travelers, particularly high-occupancy public transportation, rather than focusing solely on vehicle-centric performance. This paper proposes STDSH-MARL (Spatio-Temporal Dual-Stage Hypergraph based Multi-Agent Reinforcement Learning), a multi-agent deep reinforcement learning framework that follows a centralized training and decentralized execution paradigm. The proposed method captures spatio-temporal dependencies through a novel dual-stage hypergraph attention mechanism that models interactions across both spatial and temporal hyperedges. In addition, a hybrid discrete action space is introduced to jointly determine the next signal phase configuration and its corresponding green duration, enabling more adaptive signal timing decisions. Experiments conducted on a corridor network under five traffic scenarios demonstrate that STDSH-MARL achieves strong overall multimodal performance, with substantial and relatively consistent reductions in tram waiting time, while improvements in bus waiting time are more variable across traffic scenarios. These results highlight the trade-off among overall network efficiency, tram priority, and bus service quality. Compared with state-of-the-art baseline methods, the proposed approach achieves superior overall performance. Further ablation studies confirm the contribution of each component of STDSH-MARL, with temporal hyperedges identified as the most influential factor driving the observed performance gains.

交通信号强化学习多智能体超图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。