对比多种强化学习控灯方法在突发交通事件下的鲁棒性表现
Robustness of Reinforcement Learning-Based Traffic Signal Control under Incidents: A Comparative Study
- 构建基于SUMO的T-REX仿真框架,模拟真实路网中车辆的动态避让与拥堵传播
- 发现独立型方法在稳定路况下快且准,但遇突发事件时性能骤降
- 分层协同方法虽训练慢,但在复杂路网中更稳定,适合实际城市部署
基于强化学习的交通信号控制(RL-TSC)在提升城市交通效率方面展现出巨大潜力,但其在真实世界突发事件(如交通事故)下的鲁棒性仍缺乏系统研究。本文提出T-REX——一个基于SUMO的开源仿真框架,用于在动态、事件驱动场景下训练与评估RL-TSC方法。该框架综合考虑驾驶员概率性绕行、速度调节及上下文相关变道行为,可真实模拟事件引发的拥堵传播。为评估鲁棒性,我们设计了一套超越传统效率指标的评测体系。通过在合成与真实路网中的大量实验,验证了多个前沿RL-TSC方法在不同部署模式下的表现。结果表明:独立值函数与去中心化压力型方法在平稳交通与同质网络中收敛快、泛化好,但在事件导致的分布偏移下性能显著下降;而分层协调方法在大规模、非规则路网中表现出更强的稳定性与适应性,得益于其结构化的决策架构,但伴随训练复杂度高、收敛慢的代价。研究强调了在RL-TSC研究中需重视鲁棒性设计与评估。T-REX为此提供了开放、标准、可复现的基准平台,支持在动态扰动场景下对各类方法进行公平评测。
原文摘要 · Abstract (English)
Reinforcement learning-based traffic signal control (RL-TSC) has emerged as a promising approach for improving urban mobility. However, its robustness under real-world disruptions such as traffic incidents remains largely underexplored. In this study, we introduce T-REX, an open-source, SUMO-based simulation framework for training and evaluating RL-TSC methods under dynamic, incident scenarios. T-REX models realistic network-level performance considering drivers' probabilistic rerouting, speed adaptation, and contextual lane-changing, enabling the simulation of congestion propagation under incidents. To assess robustness, we propose a suite of metrics that extend beyond conventional traffic efficiency measures. Through extensive experiments across synthetic and real-world networks, we showcase T-REX for the evaluation of several state-of-the-art RL-TSC methods under multiple real-world deployment paradigms. Our findings show that while independent value-based and decentralized pressure-based methods offer fast convergence and generalization in stable traffic conditions and homogeneous networks, their performance degrades sharply under incident-driven distribution shifts. In contrast, hierarchical coordination methods tend to offer more stable and adaptable performance in large-scale, irregular networks, benefiting from their structured decision-making architecture. However, this comes with the trade-off of slower convergence and higher training complexity. These findings highlight the need for robustness-aware design and evaluation in RL-TSC research. T-REX contributes to this effort by providing an open, standardized and reproducible platform for benchmarking RL methods under dynamic and disruptive traffic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。