用视觉语言模型指导自动驾驶,让机器学会人类的驾驶逻辑。
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
- 以视觉语言模型为教师,注入常识推理与动作标签进行训练
- 在nuScenes上提升规划准确率,碰撞率显著降低
- 无需在部署时使用大模型,适合实时自动驾驶系统
人类驾驶员依赖常识推理应对复杂多变的真实场景。现有端到端自动驾驶模型通常仅模仿数据中的驾驶模式,未能捕捉背后的决策逻辑,限制了其处理挑战性场景的能力。为此,我们提出VLM-AD,利用视觉语言模型(VLMs)作为教师,提供包含非结构化推理信息和结构化动作标签的额外监督信号,增强模型对驾驶行为背后原因的理解能力。该方法在推理阶段不依赖VLM,具备实际部署可行性。集成至当前最优方法后,VLM-AD在nuScenes数据集上显著提升规划精度、降低碰撞率;在闭环评估中进一步提高路线完成率与驾驶评分,验证了其在长时序交互驾驶场景中的有效性,具备安全可靠落地潜力。
原文摘要 · Abstract (English)
Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing the underlying reasoning processes. This limitation constrains their ability to handle challenging driving scenarios. To close this gap, we propose VLM-AD, a method that leverages vision-language models (VLMs) as teachers to enhance training by providing additional supervision that incorporates unstructured reasoning information and structured action labels. Such supervision enhances the model's ability to learn richer feature representations that capture the rationale behind driving patterns. Importantly, our method does not require a VLM during inference, making it practical for real-time deployment. When integrated with state-of-the-art methods, VLM-AD achieves significant improvements in planning accuracy and reduced collision rates on the nuScenes dataset. It further improves route completion and driving scores under closed-loop evaluation, demonstrating its effectiveness in long-horizon, interactive driving scenarios and its potential for safe and reliable real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。