arXiv:2604.16406cs.AIcs.LG2026-04被引 2

用自洽生成框架模拟高速路复杂交通,零样本适配真实高交互场景。

Heterogeneous Self-Play for Realistic Highway Traffic Simulation

论文配图:Heterogeneous Self-Play for Realistic Highway Traffic Simulation
图 1 · 摘自论文原文
  • 按车辆类型分层建模,通过上下文条件动作实现多智能体可控交互
  • 合成场景覆盖全速域,零样本迁移在512个真实场景中成功率96.3%
  • 无需专家数据,生成轨迹更贴近真实行为,距离指标提升超13%

真实高速交通模拟对自动驾驶安全评估至关重要,尤其针对日志数据难以覆盖的稀有交互场景。然而,高速交通生成仍面临三大挑战:速度与动作覆盖广度、罕见危急场景可控生成、多智能体交互行为可信度。本文提出PHASE(Policy for Heterogeneous Agent Self-play on Expressway),一种上下文感知的自洽学习框架,通过显式单智能体条件控制实现可调性,合成场景生成保证高速覆盖,闭环多智能体训练保障交互动态真实。该框架支持单一策略下乘客车与挂车等异构车辆类型,基于车辆感知动力学与上下文条件动作实现统一建模,并通过不可恢复状态提前终止、责任碰撞判定、高速奖励塑形、耦合课程学习与鲁棒策略优化稳定自洽训练。尽管仅在合成数据上训练,PHASE在exiD的512个未见高交互真实场景中实现96.3%成功率,将ADE/FDE从6.57/12.07米降至2.44/5.25米,优于先前自洽基线。在学习轨迹嵌入空间中,相比IDM模型,弗雷切特轨迹距离降低13.1%,能量距离下降20.2%。结果表明,合成自洽训练可实现无需直接模仿专家日志的可控且真实高速场景生成。

原文摘要 · Abstract (English)

Realistic highway simulation is critical for scalable safety evaluation of autonomous vehicles, particularly for interactions that are too rare to study from logged data alone. Yet highway traffic generation remains challenging because it requires broad coverage across speeds and maneuvers, controllable generation of rare safety-critical scenarios, and behavioral credibility in multi-agent interactions. We present PHASE, Policy for Heterogeneous Agent Self-play on Expressway, a context-aware self-play framework that addresses these three requirements through explicit per-agent conditioning for controllability, synthetic scenario generation for broad highway coverage, and closed-loop multi-agent training for realistic interaction dynamics. PHASE further supports different vehicle profiles, for example, passenger cars and articulated trailer trucks, within a single policy via vehicle-aware dynamics and context-conditioned actions, and stabilizes self-play with early termination of unrecoverable states, at-fault collision attribution, highway-aware reward shaping, coupled curricula, and robust policy optimization. Despite being trained only on synthetic data, PHASE transfers zero-shot to 512 unseen high-interaction real scenarios in exiD, achieving a 96.3% success rate and reducing ADE/FDE from 6.57/12.07 m to 2.44/5.25 m relative to a prior self-play baseline. In a learned trajectory embedding space, it also improves behavioral realism over IDM, reducing Frechet trajectory distance by 13.1% and energy distance by 20.2%. These results show that synthetic self-play can provide a scalable route to controllable and realistic highway scenario generation without direct imitation of expert logs.

交通仿真自洽学习多智能体生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。