arXiv:2503.11488cs.LGcs.AI2025-03被引 5

提出通用协作强化学习框架,提升复杂交通网信号控制的适应性与泛化能力。

Unicorn: A Universal and Collaborative Reinforcement Learning Approach Towards Generalizable Network-Wide Traffic Signal Control

  • 统一异构路口状态动作表示,基于交通流映射到共同结构
  • 设计解码器网络提取通用特征,结合变分推断捕捉路口特异性
  • 自监督对比学习增强特征区分,支持区域协同优化

自适应交通信号控制(ATSC)在快速发展的城市中对缓解拥堵、提升通行效率和改善出行体验至关重要。尽管参数共享的多智能体强化学习(MARL)已在大规模同质网络中显著提升了复杂动态流的可扩展优化能力,但真实交通网络中存在的拓扑差异和交互动态的异质性,仍严重制约了ATSC在不同场景下的可扩展性与有效性。为此,本文提出Unicorn——一个面向高效、可适应的全域交通信号控制的通用协作式MARL框架。首先,提出一种统一方法,将具有不同拓扑结构的路口状态与动作映射至统一的交通流结构;其次,设计仅含解码器的通用交通表征(UTR)模块,用于提取跨场景通用特征;同时引入路口特性表征(ISR)模块,通过变分推断识别代表路口拓扑与动态的关键潜在向量;为进一步优化潜在表征,采用自监督对比学习方法,增强对路口特异性特征的区分能力;此外,将邻近智能体的状态-动作依赖关系融入策略优化过程,有效捕捉动态交互,促进区域协同。代码已开源。

原文摘要 · Abstract (English)

Adaptive traffic signal control (ATSC) is crucial in reducing congestion, maximizing throughput, and improving mobility in rapidly growing urban areas. Recent advancements in parameter-sharing multi-agent reinforcement learning (MARL) have greatly enhanced the scalable and adaptive optimization of complex, dynamic flows in large-scale homogeneous networks. However, the inherent heterogeneity of real-world traffic networks, with their varied intersection topologies and interaction dynamics, poses substantial challenges to achieving scalable and effective ATSC across different traffic scenarios. To address these challenges, we present Unicorn, a universal and collaborative MARL framework designed for efficient and adaptable network-wide ATSC. Specifically, we first propose a unified approach to map the states and actions of intersections with varying topologies into a common structure based on traffic movements. Next, we design a Universal Traffic Representation (UTR) module with a decoder-only network for general feature extraction, enhancing the model's adaptability to diverse traffic scenarios. Additionally, we incorporate an Intersection Specifics Representation (ISR) module, designed to identify key latent vectors that represent the unique intersection's topology and traffic dynamics through variational inference techniques. To further refine these latent representations, we employ a contrastive learning approach in a self-supervised manner, which enables better differentiation of intersection-specific features. Moreover, we integrate the state-action dependencies of neighboring agents into policy optimization, which effectively captures dynamic agent interactions and facilitates efficient regional collaboration. [...]. The code is available at https://github.com/marmotlab/Unicorn

交通控制强化学习多智能体泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。