让视觉语言模型与端到端驾驶策略对齐,提升决策与规划一致性。
Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning
- 通过三阶段训练显式对齐视觉语言模型与端到端策略
- 开放环下决策一致性提升19.3% F1,闭合环安全性能提升30.6%
- 适合关注自动驾驶系统协同设计的研究者与工程师
视觉语言模型(VLM)通过高层次语义推理增强了端到端(E2E)驾驶策略的规划能力。然而,现有方法常忽略VLM高层决策与E2E低层规划之间的双系统一致性问题,导致生成轨迹与意图不符,削弱了系统的自上而下引导能力。为此,我们提出Senna-2,一种显式对齐双系统的先进VLM-E2E驾驶策略。方法采用一致性导向的三阶段训练范式:第一阶段进行驾驶预训练,通过决策适配器以隐式嵌入形式传递VLM决策;第二阶段在开环设置下对齐VLM与E2E策略;第三阶段在3DGS环境中通过自下而上的分层强化学习实现闭环对齐,提升安全性和效率。大量实验表明,Senna-2在双系统一致性上取得显著提升(F1分数提高19.3%),并显著改善驾驶安全性:开环下最终位移误差降低5.7%,闭环下异常碰撞率下降30.6%。
原文摘要 · Abstract (English)
Vision-language models (VLMs) enhance the planning capability of end-to-end (E2E) driving policy by leveraging high-level semantic reasoning. However, existing approaches often overlook the dual-system consistency between VLM's high-level decision and E2E's low-level planning. As a result, the generated trajectories may misalign with the intended driving decisions, leading to weakened top-down guidance and decision-following ability of the system. To address this issue, we propose Senna-2, an advanced VLM-E2E driving policy that explicitly aligns the two systems for consistent decision-making and planning. Our method follows a consistency-oriented three-stage training paradigm. In the first stage, we conduct driving pre-training to achieve preliminary decision-making and planning, with a decision adapter transmitting VLM decisions to E2E policy in the form of implicit embeddings. In the second stage, we align the VLM and the E2E policy in an open-loop setting. In the third stage, we perform closed-loop alignment via bottom-up Hierarchical Reinforcement Learning in 3DGS environments to reinforce the safety and efficiency. Extensive experiments demonstrate that Senna-2 achieves superior dual-system consistency (19.3% F1 score improvement) and significantly enhances driving safety in both open-loop (5.7% FDE reduction) and closed-loop settings (30.6% AF-CR reduction).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。