用视觉语言模型提升自动驾驶决策,让系统更安全透明。
VLAD: A VLM-Augmented Autonomous Driving Framework with Hierarchical Planning and Interpretable Decision Process
- 用定制问答数据微调VLM,增强空间推理能力
- 在nuScenes数据集上碰撞率降低31.82%
- 生成自然语言解释,提升黑箱系统的可解释性
近期开源视觉语言模型(如LLaVA、Qwen-VL、Llama)的发展推动了其在各类系统中的集成研究。这些模型所蕴含的互联网级通用知识为提升自动驾驶的感知、预测与规划能力提供了巨大机遇。本文提出VLAD,一种融合微调后视觉语言模型(VLM)与端到端系统VAD的自动驾驶框架。通过定制问答数据集对VLM进行专项微调,强化其空间推理能力,生成高层导航指令供VAD处理以控制车辆。此外,系统还能输出可解释的自然语言决策说明,提升传统端到端架构的透明度与可信度。在真实世界nuScenes数据集上的综合评估表明,该集成系统相较基线方法平均碰撞率降低31.82%,确立了新型VLM增强型自动驾驶系统的新基准。
原文摘要 · Abstract (English)
Recent advancements in open-source Visual Language Models (VLMs) such as LLaVA, Qwen-VL, and Llama have catalyzed extensive research on their integration with diverse systems. The internet-scale general knowledge encapsulated within these models presents significant opportunities for enhancing autonomous driving perception, prediction, and planning capabilities. In this paper we propose VLAD, a vision-language autonomous driving model, which integrates a fine-tuned VLM with VAD, a state-of-the-art end-to-end system. We implement a specialized fine-tuning approach using custom question-answer datasets designed specifically to improve the spatial reasoning capabilities of the model. The enhanced VLM generates high-level navigational commands that VAD subsequently processes to guide vehicle operation. Additionally, our system produces interpretable natural language explanations of driving decisions, thereby increasing transparency and trustworthiness of the traditionally black-box end-to-end architecture. Comprehensive evaluation on the real-world nuScenes dataset demonstrates that our integrated system reduces average collision rates by 31.82% compared to baseline methodologies, establishing a new benchmark for VLM-augmented autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。