融合认知推理与端到端驾驶,提升长尾场景泛化能力
DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
- 用视觉语言动作模块生成语义驱动轨迹
- 通过密集轨迹词汇保证物理可行性,提升49.12的EPDMS得分
- 以度量引导评分器对齐双模块输出,融合互补优势
传统端到端驾驶模型虽能生成合理轨迹,但在长尾场景下因缺乏世界知识而泛化能力差。相比之下,视觉-语言-动作(VLA)模型具备世界知识,但3D推理能力有限,易生成不合理的动作。本文提出DiffVLA++框架,通过度量引导对齐机制,显式连接认知推理与端到端规划。首先构建直接生成语义驱动轨迹的VLA模块;其次设计具有密集轨迹词汇的端到端模块,确保物理可行性;最关键的是引入度量引导轨迹评分器,协调两模块输出,融合各自优势。在ICCV 2025自动驾驶挑战赛排行榜上的实验表明,DiffVLA++取得49.12的EPDMS得分。
原文摘要 · Abstract (English)
Conventional end-to-end (E2E) driving models are effective at generating physically plausible trajectories, but often fail to generalize to long-tail scenarios due to the lack of essential world knowledge to understand and reason about surrounding environments. In contrast, Vision-Language-Action (VLA) models leverage world knowledge to handle challenging cases, but their limited 3D reasoning capability can lead to physically infeasible actions. In this work we introduce DiffVLA++, an enhanced autonomous driving framework that explicitly bridges cognitive reasoning and E2E planning through metric-guided alignment. First, we build a VLA module directly generating semantically grounded driving trajectories. Second, we design an E2E module with a dense trajectory vocabulary that ensures physical feasibility. Third, and most critically, we introduce a metric-guided trajectory scorer that guides and aligns the outputs of the VLA and E2E modules, thereby integrating their complementary strengths. The experiment on the ICCV 2025 Autonomous Grand Challenge leaderboard shows that DiffVLA++ achieves EPDMS of 49.12.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。