探究轻量级视觉语言模型在自动驾驶中的视觉概念表征能力
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
- 通过对抗性图像集与线性探测,分析视觉概念在模型激活中的编码方式
- 发现物体存在可线性编码,但方向等空间概念多为隐式编码或未编码
- 识别出感知与认知两类失败模式,适用于自动驾驶模型可解释性研究
视觉语言模型(VLMs)在自动驾驶中的应用日益广泛,旨在利用其推理与泛化能力应对长尾场景。然而,这些模型在处理高度相关的简单视觉问题时仍频繁失败,且失败原因尚不明确。本文通过分析VLM中间激活,评估特定视觉概念在线性空间中的编码程度,以识别视觉信息流的瓶颈。具体而言,构建仅在目标视觉概念上不同的对抗性图像集,并使用五种前沿(SOTA)VLMs(包括一种模型的两个训练变体)的激活进行线性探测。结果表明,如物体或主体的存在等概念被显式且线性编码,而对象或主体的方向等空间概念则仅通过视觉编码器保留的空间结构隐式编码,或根本未线性编码。同时,我们观察到即使某概念在线性空间中被编码,模型仍可能错误回答。据此,我们识别出两种失败模式:感知失败(所需视觉信息未线性编码)和认知失败(视觉信息存在但与语言语义对齐失败)。最后,我们发现当目标物体距离增大时,对应视觉概念的线性可分性迅速下降。
原文摘要 · Abstract (English)
The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning and generalisation capabilities to handle long-tail scenarios. However, these models often fail on simple visual questions that are highly relevant to automated driving, and the reasons behind these failures remain poorly understood. In this work, we examine the intermediate activations of VLMs and assess the extent to which specific visual concepts are linearly encoded, with the goal of identifying bottlenecks in the flow of visual information. Specifically, we create counterfactual image sets that differ only in a targeted visual concept and then train linear probes to distinguish between them using the activations of five state-of-the-art (SOTA) VLMs, including two training variants for one model. Our results show that concepts such as the presence of an object or agent in a scene are explicitly and linearly encoded, whereas other spatial visual concepts, such as the orientation of an object or agent, are only implicitly encoded by the spatial structure retained by the vision encoder, or are not linearly encoded at all. In parallel, we observe that in certain cases, even when a concept is linearly encoded in the model's activations, the model still fails to answer correctly. This leads us to identify two failure modes. The first is perceptual failure, where the visual information required to answer a question is not linearly encoded in the model's activations. The second is cognitive failure, where the visual information is present but the model fails to align it correctly with language semantics. Finally, we show that increasing the distance of the object in question quickly degrades the linear separability of the corresponding visual concept...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。