用强化学习动态调信号灯,让城市路口更畅通
Reinforcement Learning-Based Traffic Signal Control for IoT-Enabled Intersections
- 基于PPO算法的边缘智能控制器,根据实时车流自适应分配绿灯时长
- 相比固定时序减少46%车辆等待时间,减排23%二氧化碳
- 无需中心调度、抗干扰强,适合智慧交通系统落地
城市交通拥堵仍是依赖汽车的城市持续存在的难题,带来巨大经济与社会成本。交通信号系统正作为智慧城市基础设施中的网络化信息物理组件广泛应用,分布式感知与边缘智能使自适应交通管理成为可能。本文研究了强化学习(RL)作为一种边缘智能方法,在科威特一个信号控制路口的自适应信号运行中的应用。开发了一种基于近端策略优化(PPO)的控制器,利用本地观测到的交通状态动态分配绿灯时长,无需依赖未来需求信息或集中式协调。该控制器在基于科威特真实每小时交通量数据构建的现实仿真环境中进行评估,并与传统的固定时序控制和代表当前实践的车辆感应式控制器进行对比,以平均车辆延迟、队列长度和排放为性能指标。在正常条件下,所提控制器相比固定时序控制降低平均车辆延迟46%,相比感应控制降低34%,同时每辆车的二氧化碳排放降低约23%。这些性能提升在±15%需求扰动下依然保持,可推广至工作日与周末交通模式,且经奖励函数消融实验验证;五次随机种子下方差低,证明其统计可靠性。这些发现表明,基于学习的边缘交通信号控制是物联网赋能智慧交通系统的可行基础,也是迈向完全互联的车联网(IoV)城市出行的可部署前奏。
原文摘要 · Abstract (English)
Urban traffic congestion remains a persistent challenge in car-dependent cities, imposing significant economic and societal costs. Traffic signal systems are increasingly deployed as networked cyber-physical components within smart-city infrastructures, where distributed sensing and edge intelligence enable adaptive traffic management. This paper investigates reinforcement learning (RL) as an edge-intelligent approach for adaptive traffic signal operation at a signalized urban intersection in Kuwait. A Proximal Policy Optimization (PPO)-based controller is developed to dynamically allocate green-phase durations using locally observed traffic states, without relying on future demand information or centralized coordination. The controller is evaluated in a realistic simulation environment informed by real-world hourly traffic volume data from Kuwait, and is compared against both conventional fixed-time control and a vehicle-actuated controller representing the current state of practice, using average vehicle delay, queue length, and emissions as performance metrics. Under nominal conditions, the proposed controller reduces average vehicle delay by 46% relative to fixed-time control and 34% relative to actuated control, while also lowering per-vehicle CO2 emissions by approximately 23%. These performance gains persist under demand perturbations of +/-15%, generalize from weekday to weekend traffic patterns, and are corroborated by a reward function ablation; low variance across five random seeds confirms their statistical reliability. These findings demonstrate the practicality of learning-based edge traffic signal control as a building block for IoT-enabled smart-city transportation systems, and as a deployable precursor toward fully connected, Internet of Vehicles (IoV)-based urban mobility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。