arXiv:2602.23896cs.RO2026-02

TSC通过动态优先图实现无通信下的多车协同驾驶,提升密集交通安全性。

TSC: Topology-Conditioned Stackelberg Coordination for Multi-Agent Reinforcement Learning in Interactive Driving

  • 基于轨迹编织关系构建动态优先图,定义局部领导者-跟随者关系。
  • 在四种密集场景中碰撞率显著降低,交通效率与控制平稳性保持领先。
  • 适合高密度交互场景的自动驾驶系统,尤其注重安全性的部署。

在高密度交通中实现安全高效的自动驾驶本质上是去中心化的多智能体协调问题,冲突点如并线和交织处的交互必须在部分可观测条件下可靠解决。仅依赖局部不完整信息时,交互模式变化迅速,常导致振荡让行或不安全承诺等不稳定行为。现有MARL方法或采用同步决策,加剧非平稳性;或依赖集中式排序机制,随交通密度增加而扩展性差。为此,我们提出拓扑条件化的斯塔克尔伯格协调(TSC),一种无需通信的去中心化交互驾驶学习框架。该框架从轨迹间的类辫状交织关系中提取时变有向优先图,从而在不构建全局行动顺序的前提下定义局部领导-跟随依赖。在此图约束下,TSC将密集交互分解为图内局部斯塔克尔伯格子博弈,并在中央训练、去中心化执行(CTDE)范式下,通过动作预测预判领导者,以动作条件价值学习训练跟随者逼近局部最优响应,提升训练稳定性与密集交通下的安全性。四个密集交通场景的实验表明,TSC在关键指标上优于代表性MARL基线,尤其显著减少碰撞,同时保持良好的交通效率与控制平滑性。

原文摘要 · Abstract (English)

Safe and efficient autonomous driving in dense traffic is fundamentally a decentralized multi-agent coordination problem, where interactions at conflict points such as merging and weaving must be resolved reliably under partial observability. With only local and incomplete cues, interaction patterns can change rapidly, often causing unstable behaviors such as oscillatory yielding or unsafe commitments. Existing multi-agent reinforcement learning (MARL) approaches either adopt synchronous decision-making, which exacerbate non-stationarity, or depend on centralized sequencing mechanisms that scale poorly as traffic density increases. To address these limitations, we propose Topology-conditioned Stackelberg Coordination (TSC), a learning framework for decentralized interactive driving under communication-free execution, which extracts a time-varying directed priority graph from braid-inspired weaving relations between trajectories, thereby defining local leader-follower dependencies without constructing a global order of play. Conditioned on this graph, TSC endogenously factorizes dense interactions into graph-local Stackelberg subgames and, under centralized training and decentralized execution (CTDE), learns a sequential coordination policy that anticipates leaders via action prediction and trains followers through action-conditioned value learning to approximate local best responses, improving training stability and safety in dense traffic. Experiments across four dense traffic scenarios show that TSC achieves superior performance over representative MARL baselines across key metrics, most notably reducing collisions while maintaining competitive traffic efficiency and control smoothness.

多智能体强化学习自动驾驶协同控制交通仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。