arXiv:2605.14201cs.ROcs.CV2026-05被引 1

让自动驾驶模型在潜在空间中实时互动演练,提升闭环行驶稳定性。

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

论文配图:MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
图 1 · 摘自论文原文
  • 在视觉语言动作模型潜空间中实现多车协同闭环推演
  • 在Bench2Drive上达到当前最优驾驶性能,支持多步前瞻决策
  • 无需外部仿真器,可扩展性强,适合真实场景训练

视觉-语言-动作(VLA)模型作为端到端运动规划器有效,但在闭环评估中易因传统模仿学习训练而表现脆弱。现有闭环监督方法缺乏可扩展性,无法完整建模动态交互环境。本文提出MAPLE框架,在VLA模型的潜空间中实现动态驾驶场景下的反应式多智能体推演。自车与周边交通参与者独立控制,多步时域内对场景中其他代理做出反应,支持闭环训练。MAPLE包含两个阶段:(1) 基于真实轨迹的潜空间监督微调;(2) 结合全局与个体奖励的强化学习,鼓励安全、前进与交互真实性。进一步引入多样性奖励,促使模型生成数据中未见的规划行为。显著地,本框架可扩展且无需外部仿真器,避免高计算开销和低视觉保真度问题。MAPLE在Bench2Drive上取得最先进驾驶表现,实现可扩展的闭环多智能体交互,推动鲁棒端到端自动驾驶系统发展。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional imitation learning framework. Existing closed-loop supervision approaches lack scalability and fail to completely model a reactive environment. We propose MAPLE, a novel framework for reactive, multi-agent rollout of a dynamic driving scenario in the latent space of the VLA model. The ego vehicle and nearby traffic agents are independently controlled over multi-step horizons, while being reactive to other agents in the scene, enabling closed-loop training. MAPLE consists of two training stages: (1) supervised fine-tuning on the latent rollouts based on ground-truth trajectories, followed by (2) reinforcement learning with global and agent -specific rewards that encourage safety, progress, and interaction realism. We further propose diversity rewards that encourage the model to generate planning behaviors that may not be present in logged driving data. Notably, our closed-loop training framework is scalable and does not require external simulators, which can be computationally expensive to run and have limited visual fidelity to the real-world. MAPLE achieves state-of-the-art driving performance on Bench2Drive and demonstrates scalable, closed-loop multi-agent play for robust E2E autonomous driving systems.

自动驾驶多智能体端到端强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。