arXiv:2501.04003cs.CVcs.RO2025-01ICCV被引 136

测试视觉语言模型在自动驾驶中的可靠性,发现其常依赖文本而非真实视觉信息。

Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

  • 构建17种场景的DriveBench数据集,评估12个VLM在多种输入下的表现
  • 80%以上响应基于文本或通用知识,非真实视觉感知,尤其在图像损坏时
  • 提出新评估指标,强调多模态理解与抗干扰能力,适合安全关键系统研究

视觉语言模型(VLMs)在自动驾驶中生成可解释决策的潜力引发关注,但其是否具备真正的视觉基础、可靠性和可解释性尚未被充分验证。为此,我们构建了DriveBench基准数据集,涵盖17种设置(干净、受损、纯文本输入),包含19,200帧图像、20,498个问答对、三种问题类型、四种主流驾驶任务,以及12个主流VLM。结果表明,多数情况下,VLM生成的回答依赖于通用知识或文本线索,而非真实视觉信息,尤其在视觉输入缺失或退化时表现更差。这种现象被数据集偏差和不足的评估指标掩盖,在自动驾驶等高风险场景中存在严重隐患。此外,VLM在多模态推理上表现薄弱,对输入噪声敏感,性能波动明显。为此,我们提出强化视觉基础与多模态理解的新评估指标,并指出利用模型对输入损坏的感知能力可提升可靠性,为构建可信、可解释的自动驾驶决策系统提供路径。该基准工具包已公开。

原文摘要 · Abstract (English)

Recent advancements in Vision-Language Models (VLMs) have sparked interest in their use for autonomous driving, particularly in generating interpretable driving decisions through natural language. However, the assumption that VLMs inherently provide visually grounded, reliable, and interpretable explanations for driving remains largely unexamined. To address this gap, we introduce DriveBench, a benchmark dataset designed to evaluate VLM reliability across 17 settings (clean, corrupted, and text-only inputs), encompassing 19,200 frames, 20,498 question-answer pairs, three question types, four mainstream driving tasks, and a total of 12 popular VLMs. Our findings reveal that VLMs often generate plausible responses derived from general knowledge or textual cues rather than true visual grounding, especially under degraded or missing visual inputs. This behavior, concealed by dataset imbalances and insufficient evaluation metrics, poses significant risks in safety-critical scenarios like autonomous driving. We further observe that VLMs struggle with multi-modal reasoning and display heightened sensitivity to input corruptions, leading to inconsistencies in performance. To address these challenges, we propose refined evaluation metrics that prioritize robust visual grounding and multi-modal understanding. Additionally, we highlight the potential of leveraging VLMs' awareness of corruptions to enhance their reliability, offering a roadmap for developing more trustworthy and interpretable decision-making systems in real-world autonomous driving contexts. The benchmark toolkit is publicly accessible.

视觉语言模型自动驾驶可靠性评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。