arXiv:2603.12096cs.AI2026-03被引 1

提出新框架提升交通信号控制的鲁棒性与效率

A Robust and Efficient Multi-Agent Reinforcement Learning Framework for Traffic Signal Control

  • 引入转向比例随机化增强模型对动态车流的适应能力
  • 通过指数相位调整降低平均等待时间超10%
  • 适合需要高稳定性和泛化能力的智能交通系统

交通信号控制中的强化学习面临真实场景部署难题,主要源于对动态交通流变化的泛化能力不足。现有方法常过拟合静态模式,且动作空间与驾驶员预期不匹配。本文提出一种鲁棒的多智能体强化学习框架,在Vissim仿真器中验证。该框架包含三项机制:(1) 转向比例随机化训练策略,使智能体暴露于动态转向概率中,提升对未见场景的鲁棒性;(2) 面向稳定性的指数相位时长调整动作空间,通过周期性指数调节实现响应与精度的平衡;(3) 基于邻居观测的观察方案,结合MAPPO算法与集中式训练、分布式执行(CTDE)范式。利用集中更新近似全局观测效果,同时保持可扩展的本地通信。实验表明,该框架优于标准强化学习基线,平均等待时间减少超过10%,在未见交通场景中展现优异泛化能力,并维持高控制稳定性,为自适应信号控制提供实用解决方案。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) in Traffic Signal Control (TSC) faces significant hurdles in real-world deployment due to limited generalization to dynamic traffic flow variations. Existing approaches often overfit static patterns and use action spaces incompatible with driver expectations. This paper proposes a robust Multi-Agent Reinforcement Learning (MARL) framework validated in the Vissim traffic simulator. The framework integrates three mechanisms: (1) Turning Ratio Randomization, a training strategy that exposes agents to dynamic turning probabilities to enhance robustness against unseen scenarios; (2) a stability-oriented Exponential Phase Duration Adjustment action space, which balances responsiveness and precision through cyclical, exponential phase adjustments; and (3) a Neighbor-Based Observation scheme utilizing the MAPPO algorithm with Centralized Training with Decentralized Execution (CTDE). By leveraging centralized updates, this approach approximates the efficacy of global observations while maintaining scalable local communication. Experimental results demonstrate that our framework outperforms standard RL baselines, reducing average waiting time by over 10%. The proposed model exhibits superior generalization in unseen traffic scenarios and maintains high control stability, offering a practical solution for adaptive signal control.

交通控制强化学习多智能体鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。