arXiv:2605.09153cs.ROcs.AI2026-05

提出分层框架,让交通模拟更像真实人类驾驶。

Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation

论文配图:Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation
图 1 · 摘自论文原文
  • 高层用博弈策略生成交互意图,底层生成连续运动轨迹。
  • 在城市路网中实现更平滑安全的控制,效率媲美自对弈方法。
  • 适合需要真实驾驶行为的交通仿真与自动驾驶测试。

闭环交通仿真需要既可扩展又行为真实的智能体。近期自对弈强化学习方法虽具强可扩展性,但其均衡策略无法捕捉真实人类驾驶员的社会性行为。本文提出一种分层架构,超越自对弈:高层采用斯塔克尔伯格式多智能体强化学习(MARL)模块生成交互感知的意图指令;低层连续运动模块根据这些指令生成物理一致、场景响应的控制序列。为缓解闭环部署中的分布偏移,引入混合协同训练方案,结合MARL与辅助恢复监督。在基于SUMO的城市网络实验表明,该框架在控制平滑性和安全性上优于自对弈与被动模仿基线,同时保持竞争力的交通效率。

原文摘要 · Abstract (English)

Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches demonstrate strong scalability, but their equilibrium strategies fail to capture the socially aware behaviors of real human drivers. We propose a hierarchical architecture that goes beyond self-play by combining high-level multi-agent interaction reasoning with low-level continuous trajectory realization. Specifically, a Stackelberg-style Multi-Agent Reinforcement Learning (MARL) module generates interaction-aware intention commands. These commands condition a low-level continuous motion module, translating the strategic intent into physically consistent, scene-responsive control sequences. To mitigate distribution shift in closed-loop deployment, we introduce a hybrid co-training scheme combining MARL with auxiliary recovery supervision. Experiments on a SUMO-based urban network demonstrate that the proposed framework achieves superior control smoothness and safety compared to self-play and passive imitation baselines, while maintaining competitive traffic efficiency.

交通仿真强化学习分层决策自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。