提出BetterCheck检测视觉语言模型在自动驾驶感知中的幻觉问题。
BetterCheck: Towards Safeguarding VLMs for Automotive Perception Systems
- 设计新方法BetterCheck,识别VLM在交通场景描述中的虚构内容。
- 在Waymo数据集上测试发现,顶级VLM仍会编造不存在的交通参与者。
- 适合关注自动驾驶安全、视觉语言模型可靠性的研究者使用。
大型语言模型(LLMs)正被拓展用于同时处理文本和视频等多模态数据。其在理解图像内容方面表现卓越,超越了仅支持有限词汇的专用神经网络(如Yolo)。当不受限制时,最先进的视觉语言模型(VLMs)能描述复杂交通状况,具备成为自动驾驶感知系统组成部分的潜力。然而,这些模型易产生幻觉:可能遗漏真实存在的弱势道路使用者,或错误地“看到”并不存在的交通参与者。前者可能导致灾难性决策,后者则引发不必要的减速。本文系统评估了三种先进VLM在从Waymo Open Dataset中采样的多样化交通场景上的表现,旨在为基于VLM的感知系统建立安全防护机制。结果显示,无论是开源还是专有模型,均展现出极强的图像理解能力,甚至能注意到人类难以察觉的细微细节。但它们仍会编造描述内容,亟需如BetterCheck所提出的幻觉检测策略来保障安全性。
原文摘要 · Abstract (English)
Large language models (LLMs) are growingly extended to process multimodal data such as text and video simultaneously. Their remarkable performance in understanding what is shown in images is surpassing specialized neural networks (NNs) such as Yolo that is supporting only a well-formed but very limited vocabulary, ie., objects that they are able to detect. When being non-restricted, LLMs and in particular state-of-the-art vision language models (VLMs) show impressive performance to describe even complex traffic situations. This is making them potentially suitable components for automotive perception systems to support the understanding of complex traffic situations or edge case situation. However, LLMs and VLMs are prone to hallucination, which mean to either potentially not seeing traffic agents such as vulnerable road users who are present in a situation, or to seeing traffic agents who are not there in reality. While the latter is unwanted making an ADAS or autonomous driving systems (ADS) to unnecessarily slow down, the former could lead to disastrous decisions from an ADS. In our work, we are systematically assessing the performance of 3 state-of-the-art VLMs on a diverse subset of traffic situations sampled from the Waymo Open Dataset to support safety guardrails for capturing such hallucinations in VLM-supported perception systems. We observe that both, proprietary and open VLMs exhibit remarkable image understanding capabilities even paying thorough attention to fine details sometimes difficult to spot for us humans. However, they are also still prone to making up elements in their descriptions to date requiring hallucination detection strategies such as BetterCheck that we propose in our work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。