arXiv:2604.11734cs.ROcs.AI2026-04被引 2

提出SCORP模型,实现多智能体协同驾驶的场景一致与稳定在线强化学习。

SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving

论文配图:SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving
图 1 · 摘自论文原文
  • 设计场景条件化扩散架构,提升轨迹与道路环境的一致性。
  • 在WOMD数据集上安全与效率指标提升10.47%至28.26%。
  • 适合需要高安全性和协同效率的自动驾驶系统研发者。

协同驾驶是关乎安全与效率的关键任务,需协调多样且交互真实的多智能体轨迹。现有基于扩散的方法虽能从示范中捕捉多模态行为,但常缺乏场景一致性,且难以对齐闭环协同目标。因此需后训练进一步优化,但在反应式多智能体环境中实现稳定在线后训练仍具挑战。本文提出SCORP,一种具备稳定在线强化学习(RL)后训练能力的场景一致多智能体扩散规划器。预训练阶段,设计场景条件化多智能体去噪架构,结合智能体间自注意力与双路径条件机制:交叉注意力直接注入场景信息,AdaLN-Zero实现灵活稳定的条件调制,从而提升联合轨迹的场景一致性与道路依从性。后训练阶段,构建两层马尔可夫决策过程(MDP),显式融合反向去噪链与策略-环境交互。同时协同设计密集、结构良好的规划奖励与方差门控组相对策略优化(VG-GRPO),缓解闭环训练中的优势崩溃与梯度不稳定性问题。大量实验表明,SCORP在WOMD数据集上优于强开源基线,核心安全与效率指标分别提升10.47%-28.26%和1.70%-7.22%。相较于其他后训练方法,SCORP在驾驶安全与交通效率上均实现显著且持续的增益,凸显其在闭环协同驾驶中的稳定进步。

原文摘要 · Abstract (English)

Cooperative driving is a safety- and efficiency-critical task that requires the coordination of diverse, interaction-realistic multi-agent trajectories. Although existing diffusion-based methods can capture multimodal behaviors from demonstrations, they often exhibit weak scene consistency and poor alignment with closed-loop cooperative objectives. This makes post-training necessary for further improvement, yet achieving stable online post-training in reactive multi-agent environments remains challenging. In this paper, we propose SCORP, a scene-consistent multi-agent diffusion planner with stable online reinforcement learning (RL) post-training for cooperative driving. For pre-training, we develop a scene-conditioned multi-agent denoising architecture that couples inter-agent self-attention with a dual-path conditioning mechanism: cross-attention provides direct scene-information injection, while AdaLN-Zero enables additional flexible and stable conditional modulation, thereby improving the scene consistency and road adherence of joint trajectories. For post-training, we formulate a two-layer Markov decision process (MDP) that explicitly integrates the reverse denoising chain with policy-environment interaction. We further co-design dense, well-shaped planning rewards and variance-gated group-relative policy optimization (VG-GRPO) to mitigate advantage collapse and gradient instability during closed-loop training. Extensive experiments show that SCORP outperforms strong open-source baselines on WOMD, with 10.47%-28.26% and 1.70%-7.22% improvements in core safety and efficiency metrics, respectively. Moreover, compared with alternative post-training methods, SCORP delivers significant and consistent gains in both driving safety and traffic efficiency, highlighting stable and sustained advances in closed-loop cooperative driving.

多智能体协同驾驶扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。