arXiv:2608.12854cs.ROcs.AI2026-08

将语义先验与预测动态协同,提升自动驾驶规划性能。

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

论文配图:BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
图 1 · 摘自论文原文
  • 分路处理语义推理与世界建模,通过动作表示对齐
  • 在NAVSIM v1和v2上分别达89.5和89.6的最优指标
  • 适合追求高精度与长程规划的自动驾驶系统研发

自动驾驶需同时满足语义约束与未来动态预测。现有端到端方法多侧重单一方向:视觉语言动作(VLA)模型依赖视觉语言模型(VLM)先验进行语义推理,而世界动作模型(WAMs)通过生成式世界建模实现未来感知。这自然催生统一规划器的需求,但简单融合会导致注意力分配失衡,语义捷径主导共享空间并抑制预测动态。受神经科学启发,我们提出BrainWAM,一种结构化动作空间协调框架,将语义推理与世界建模转为两个专用的动作导向路径,并在紧凑动作表示层面实现对齐。进一步引入异步修正流推断策略,解耦视频与动作去噪,降低推理延迟同时保留规划相关预测上下文。BrainWAM在NAVSIM v1(89.5 PDMS)和NAVSIM v2(89.6 EPDMS)上均达到当前最优性能,显著优于仅用VLA或仅用WAM的方法,验证了其在自动驾驶系统中的实用潜力。

原文摘要 · Abstract (English)

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.

自动驾驶动作规划多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。