arXiv:2412.09951cs.CV2024-12被引 33

用视觉语言模型增强自动驾驶,减少事故并提升规划能力。

WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

  • 基于驾驶知识与规划数据联合训练,实现知识对齐的轨迹规划。
  • 知识多样性提升后,事故率显著下降,驾驶得分提高11.9%。
  • 适合研究自动驾驶决策与多场景推理的学者与工程师。

视觉语言模型(VLMs)展现出强大的通用人类知识与逻辑推理能力,推动其在高阶自动驾驶任务如场景理解与决策中的应用。然而,驾驶专业知识深度与广度对闭环自动驾驶性能的影响仍需深入探究。本文研究了基础驾驶知识对闭环轨迹规划的影响,提出专用于端到端自动驾驶的WiseAD模型,具备驾驶推理、行为解释、目标识别、风险分析、驾驶建议及跨场景轨迹规划能力。通过在驾驶知识与规划数据集上联合训练,模型实现知识对齐的轨迹规划。大量实验表明,随着驾驶知识多样性的扩展,关键事故显著减少,在Carla闭环评估中驾驶得分和路线完成率分别提升11.9%和12.4%,达到当前最优水平。此外,WiseAD在域内与域外知识评估中也表现优异。

原文摘要 · Abstract (English)

The emergence of general human knowledge and impressive logical reasoning capacity in rapidly progressed vision-language models (VLMs) have driven increasing interest in applying VLMs to high-level autonomous driving tasks, such as scene understanding and decision-making. However, an in-depth study on the relationship between knowledge proficiency, especially essential driving expertise, and closed-loop autonomous driving performance requires further exploration. In this paper, we investigate the effects of the depth and breadth of fundamental driving knowledge on closed-loop trajectory planning and introduce WiseAD, a specialized VLM tailored for end-to-end autonomous driving capable of driving reasoning, action justification, object recognition, risk analysis, driving suggestions, and trajectory planning across diverse scenarios. We employ joint training on driving knowledge and planning datasets, enabling the model to perform knowledge-aligned trajectory planning accordingly. Extensive experiments indicate that as the diversity of driving knowledge extends, critical accidents are notably reduced, contributing 11.9% and 12.4% improvements in the driving score and route completion on the Carla closed-loop evaluations, achieving state-of-the-art performance. Moreover, WiseAD also demonstrates remarkable performance in knowledge evaluations on both in-domain and out-of-domain datasets.

自动驾驶视觉语言模型知识增强轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。