让视觉语言模型更懂复杂路况,提升自动驾驶可解释性。
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
- 引入场景交互与专家适配器,增强模型对驾驶场景的理解。
- 在DriveLM和nuScenes-QA上显著提升推理性能,超越现有模型。
- 兼容现有大模型,适合需要可解释性决策的自动驾驶系统。
大型视觉语言模型(VLM)在场景理解方面展现出潜力,提升了驾驶行为的可解释性与人机交互能力。现有方法主要对车载多视角图像和场景推理文本进行微调,但通常缺乏全面的场景识别能力与强空间感知,尤其在复杂情境下表现不足。为此,我们提出专为自动驾驶设计的新框架LMAD,模拟现代端到端驾驶范式,融合全面的场景理解与任务专用结构。具体而言,在同一驾驶任务结构中引入初步场景交互与专用专家适配器,使VLM更贴合实际驾驶场景。此外,该方法可完全兼容现有VLM,并无缝集成至规划导向的驾驶系统。在DriveLM与nuScenes-QA数据集上的大量实验表明,LMAD显著提升了现有VLM在驾驶推理任务中的表现,树立了可解释自动驾驶的新标准。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view images and scene reasoning text, but this approach often lacks the holistic and nuanced scene recognition and powerful spatial awareness required for autonomous driving, especially in complex situations. To address this gap, we propose a novel vision-language framework tailored for autonomous driving, called LMAD. Our framework emulates modern end-to-end driving paradigms by incorporating comprehensive scene understanding and a task-specialized structure with VLMs. In particular, we introduce preliminary scene interaction and specialized expert adapters within the same driving task structure, which better align VLMs with autonomous driving scenarios. Furthermore, our approach is designed to be fully compatible with existing VLMs while seamlessly integrating with planning-oriented driving systems. Extensive experiments on the DriveLM and nuScenes-QA datasets demonstrate that LMAD significantly boosts the performance of existing VLMs on driving reasoning tasks,setting a new standard in explainable autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。