arXiv:2605.29693cs.LGcs.RO2026-05中稿 · IEEE International…

用动量奖励优化交通灯,减少拥堵和排放。

Momentum Based Reward Design for Low Emission Traffic Signal Control

论文配图:Momentum Based Reward Design for Low Emission Traffic Signal Control
图 1 · 摘自论文原文
  • 设计动量奖励函数,鼓励车辆持续通行而非只惩罚拥堵
  • 在SUMO仿真中,吞吐量与碳排放平衡优于传统方法
  • 适合关注低碳交通控制的研究者和城市规划者

城市交通拥堵是全球性问题,加剧通勤时间与环境污染。传统信号控制系统难以适应动态交通。自适应信号控制可在不改变道路设施的情况下提升交通效率。深度强化学习(DRL)在此任务中表现优异,但现有基于延迟或队列的奖励常导致短视或不稳定的策略。本文提出动量奖励函数(MBRF),通过鼓励车辆持续通行来改善控制效果。在SUMO(Simulation of Urban MObility)中,使用等待时间、队列长度、吞吐量和二氧化碳排放等标准指标进行评估。结果表明,该方法在吞吐量-排放权衡上优于基于延迟或队列的奖励,且学习过程更稳定,同时超越了经典控制器如最大压力法(Max Pressure)和线性排队因子(LQF)。

原文摘要 · Abstract (English)

Urban traffic congestion is a growing global issue contributing significantly to long commute times and environmental pollution. Traditional traffic signal control systems often fail to adapt to dynamic traffic conditions. Adaptive traffic signal control can improve urban traffic without changing road infrastructure. Deep Reinforcement Learning (DRL) has shown strong performance for this task, but existing delay and queue-based rewards often produce short-sighted or unstable policies. This paper proposes a Momentum-Based Reward Function (MBRF) that encourages vehicles to keep moving rather than penalizing congestion alone. The method is evaluated in SUMO (Simulation of Urban MObility) using standard traffic metrics such as waiting time, queue length, throughput, and CO2 emissions. Results show that the proposed reward produces better throughput-emission trade-offs and more stable learning behavior than delay or queue-based rewards, as well as classical controllers such as Max Pressure and LQF.

交通信号控制强化学习低碳交通仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。