用视觉语言模型增强自动驾驶决策的可解释性
VLMs Guided Interpretable Decision Making for Autonomous Driving
- 让VLM充当语义增强器而非直接决策者
- 在两个基准上达到当前最佳性能
- 适合需要可解释决策的自动驾驶系统
近期自动驾驶研究尝试在视觉问答框架中使用视觉语言模型(VLMs)直接进行驾驶决策,但这些方法依赖手工设计提示,表现不一致,限制了真实场景中的鲁棒性和泛化能力。本文评估了开源先进VLMs在高阶决策任务上的表现,发现其生成可靠、上下文感知决策的能力存在关键局限。为此,提出新方法:将VLM角色从直接决策生成转向语义增强。利用其强大的场景理解能力,为基于视觉的基准数据注入结构化、语言丰富的场景描述。在此增强表示基础上,构建多模态交互架构,融合视觉与语言特征以实现更准确的决策和可解释的文本说明。此外,设计后处理优化模块,利用VLM提升预测可靠性。在两个自动驾驶基准上的大量实验表明,该方法达到当前最优性能,为集成VLM于可靠且可解释的自动驾驶系统提供了有前景的方向。
原文摘要 · Abstract (English)
Recent advancements in autonomous driving (AD) have explored the use of vision-language models (VLMs) within visual question answering (VQA) frameworks for direct driving decision-making. However, these approaches often depend on handcrafted prompts and suffer from inconsistent performance, limiting their robustness and generalization in real-world scenarios. In this work, we evaluate state-of-the-art open-source VLMs on high-level decision-making tasks using ego-view visual inputs and identify critical limitations in their ability to deliver reliable, context-aware decisions. Motivated by these observations, we propose a new approach that shifts the role of VLMs from direct decision generators to semantic enhancers. Specifically, we leverage their strong general scene understanding to enrich existing vision-based benchmarks with structured, linguistically rich scene descriptions. Building on this enriched representation, we introduce a multi-modal interactive architecture that fuses visual and linguistic features for more accurate decision-making and interpretable textual explanations. Furthermore, we design a post-hoc refinement module that utilizes VLMs to enhance prediction reliability. Extensive experiments on two autonomous driving benchmarks demonstrate that our approach achieves state-of-the-art performance, offering a promising direction for integrating VLMs into reliable and interpretable AD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。