arXiv:2512.04653cs.MAcs.AI2025-12

提出半集中训练分散执行架构,提升多路口信号控制效率。

A Semi Centralized Training Decentralized Execution Architecture for Multi Agent Deep Reinforcement Learning in Traffic Signal Control

  • 区域化分组训练,共享参数并融合局部与全局信息
  • 在不同交通密度下均优于传统集中/分散方法
  • 适配多种模型结构,迁移性强,适合智能交通系统

多智能体强化学习(MARL)已成为应对多路口自适应信号控制的有前景范式。现有方法通常采用完全集中或完全分散设计:前者面临维度灾难和单点依赖问题,后者则受限于严重部分可观测性且缺乏显式协调,导致性能不佳。为此,本文提出一种半集中训练、分散执行(SEMI-CTDE)架构,将路网划分为紧密耦合的区域,区域内进行集中训练,通过区域参数共享及复合状态与奖励设计,同时编码本地与区域信息。该架构对不同策略主干和状态-奖励组合具有高度可迁移性。基于此,我们实现了两个目标不同的模型,并通过多视角实验验证其有效性,涵盖核心组件消融分析、规则基线与全分散基线对比。结果表明,两种模型在各种交通密度与分布下均表现显著更优,具备广泛适用性。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has emerged as a promising paradigm for adaptive traffic signal control (ATSC) of multiple intersections. Existing approaches typically follow either a fully centralized or a fully decentralized design. Fully centralized approaches suffer from the curse of dimensionality, and reliance on a single learning server, whereas purely decentralized approaches operate under severe partial observability and lack explicit coordination resulting in suboptimal performance. These limitations motivate region-based MARL, where the network is partitioned into smaller, tightly coupled intersections that form regions, and training is organized around these regions. This paper introduces a Semi-Centralized Training, Decentralized Execution (SEMI-CTDE) architecture for multi intersection ATSC. Within each region, SEMI-CTDE performs centralized training with regional parameter sharing and employs composite state and reward formulations that jointly encode local and regional information. The architecture is highly transferable across different policy backbones and state-reward instantiations. Building on this architecture, we implement two models with distinct design objectives. A multi-perspective experimental analysis of the two implemented SEMI-CTDE-based models covering ablations of the architecture's core elements including rule based and fully decentralized baselines shows that they achieve consistently superior performance and remain effective across a wide range of traffic densities and distributions.

多智能体强化学习交通控制分布式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。