HALO用分层强化学习解决大规模交通信号控制的扩展与协同难题
HALO: Hierarchical Reinforcement Learning for Large-Scale Adaptive Traffic Signal Control
- 分层架构:高层全局策略用Transformer-LSTM建模全网时空依赖,下发指导信号
- 低层局部策略基于本地观测和全局上下文执行控制,实现高效协调
- 对抗目标设定机制让局部策略超越全局目标,提升整体性能
自适应交通信号控制(ATSC)对缓解现代智慧城市的交通拥堵至关重要,当前交通基础设施正演变为包含数千个感知-控制节点的互联物联网(WoT)环境。然而,现有方法面临关键的可扩展性-协调性权衡:集中式方法虽能优化全局目标,但在城市规模下计算不可行;去中心化多智能体方法虽可扩展,却缺乏网络级一致性,导致性能不佳。本文提出HALO,一种分层强化学习框架,有效解决该权衡问题。HALO将决策分为两级:高层全局引导策略采用Transformer-LSTM编码器建模全网时空依赖,广播紧凑的引导信号;低层局部路口策略在本地观测和全局上下文条件下执行去中心化控制。为增强全局-局部目标对齐,引入对抗目标设定机制,使全局策略提出挑战性但可行的网络级目标,局部策略被训练以超越这些目标,从而促进鲁棒协作。我们在多个标准基准及新构建的大规模曼哈顿类网络(2,668个路口)上进行评估,覆盖真实交通模式,包括高峰时段转换、恶劣天气和节假日激增。结果表明,HALO在小规模基准中表现竞争力,并随网络复杂度增加而愈发占优;在所有大规模场景中均表现最佳,平均通行时间降低最多6.8%,平均延迟降低5.0%,优于现有最优方法。
原文摘要 · Abstract (English)
Adaptive traffic signal control (ATSC) is essential for mitigating urban congestion in modern smart cities, where traffic infrastructure is evolving into interconnected Web-of-Things (WoT) environments with thousands of sensing-and-control nodes. However, existing methods face a critical scalability-coordination tradeoff: centralized approaches optimize global objectives but become computationally intractable at city scale, while decentralized multi-agent methods scale efficiently yet lack network-level coherence, resulting in suboptimal performance. In this paper, we present HALO, a hierarchical reinforcement learning framework that addresses this tradeoff for large-scale ATSC. HALO decouples decision-making into two levels: a high-level global guidance policy employs Transformer-LSTM encoders to model spatio-temporal dependencies across the entire network and broadcast compact guidance signals, while low-level local intersection policies execute decentralized control conditioned on both local observations and global context. To ensure better alignment of global-local objectives, we introduce an adversarial goal-setting mechanism where the global policy proposes challenging-yet-feasible network-level targets that local policies are trained to surpass, fostering robust coordination. We evaluate HALO extensively on multiple standard benchmarks, and a newly constructed large-scale Manhattan-like network with 2,668 intersections under real-world traffic patterns, including peak transitions, adverse weather and holiday surges. Results demonstrate HALO shows competitive performance and becomes increasingly dominant as network complexity grows across small-scale benchmarks, while delivering the strongest performance in all large-scale regimes, offering up to 6.8% lower average travel time and 5.0% lower average delay than the best state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。