提出双层认知框架,让自动驾驶模型能前瞻规划、主动决策。
WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

- 分语义与生成两层建模,融合3D感知与博弈推理。
- 在NAVSIM上实现92.9的SOTA PDMS得分,领先现有方法。
- 适合研究智能驾驶规划与多智能体协同的学者参考。
视觉-语言-动作(VLA)模型推动了端到端自动驾驶的发展。然而,现有方法或缺乏全面世界认知,或存在碎片化世界预见,导致模型仅能被动响应。为此,我们提出WCog-VLA,一种新型双层世界认知VLA框架,成功将语义世界预测与生成式世界演化相结合,实现主动驾驶。在语义层面,通过引入3D空间感知和代理标记,统一世界认知与推理,并支持博弈论思维链(Game-CoT)推理;在生成层面,提出对齐解耦扩散变换器(ADDT),生成符合物理规律的多智能体联合轨迹。通过场景表征对齐,ADDT显著减少去噪步骤,加速推理。为支持策略推理,我们构建了包含85,000条Game-CoT标注的大规模数据集。在NAVSIM基准测试中,WCog-VLA取得92.9的最新最优PDMS分数。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition or suffer from fragmented world foresight, inherently confining these models to reactive driving. To address this limitation, we propose WCog-VLA, a novel dual-level World-Cognitive VLA framework that successfully bridges semantic world forecasting with generative world evolution to achieve proactive autonomous driving. At the semantic level, WCog-VLA unifies world cognition and reasoning by incorporating 3D spatial perception and injecting agent tokens to capture the world dynamics, while concurrently enabling Game-theoretic Chain-of-Thought (Game-CoT) reasoning. At the generative level, we introduce the Aligned Decoupled Diffusion Transformer (ADDT) as a powerful generative world model that synthesizes physically-plausible joint multi-agent trajectories. Through scene representation alignment, ADDT reduces the number of denoising steps required and thus significantly accelerates inference. To facilitate strategic reasoning, we further construct a large-scale dataset featuring 85k Game-CoT annotations. Extensive experiments on the NAVSIM benchmark demonstrate that WCog-VLA achieves a State-Of-The-Art (SOTA) PDMS score of 92.9.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。