VLM模型在视觉任务上表现远低于直接读取视觉编码器,暴露其信息整合缺陷。
Hidden in plain sight: VLMs overlook their visual representations
- 对比视觉编码器直接读取,检验VLM多模态融合能力
- 在深度估计等任务中性能接近随机水平
- 语言模型先验导致视觉信息利用不足,适合研究多模态失效
语言为指定和评估视觉任务表现提供了自然接口。为实现这一目标,视觉语言模型(VLMs)必须有效整合视觉与语言信息。本文通过对比VLMs与其视觉编码器的直接读出结果,评估其跨模态整合能力。在一系列以视觉为中心的基准测试(如深度估计、对应关系)中,VLMs表现显著劣于其视觉编码器,性能降至接近随机水平。通过全模型分析,我们发现:1)视觉表征退化;2)对任务提示敏感;3)语言模型在任务求解中起主导作用。瓶颈在于第三类——VLM未能有效利用模型内广泛可得的视觉信息,且继承了语言模型固有的语言先验。本工作有助于诊断开源VLM的失败模式,并为未来研究提供一系列评估工具。
原文摘要 · Abstract (English)
Language provides a natural interface to specify and evaluate performance on visual tasks. To realize this possibility, vision language models (VLMs) must successfully integrate visual and linguistic information. Our work compares VLMs to a direct readout of their visual encoders to understand their ability to integrate across these modalities. Across a series of vision-centric benchmarks (e.g., depth estimation, correspondence), we find that VLMs perform substantially worse than their visual encoders, dropping to near-chance performance. We investigate these results through a series of analyses across the entire VLM: namely 1) the degradation of vision representations, 2) brittleness to task prompt, and 3) the language model's role in solving the task. We find that the bottleneck in performing these vision-centric tasks lies in this third category; VLMs are not effectively using visual information easily accessible throughout the entire model, and they inherit the language priors present in the LLM. Our work helps diagnose the failure modes of open-source VLMs, and presents a series of evaluations useful for future investigations into visual understanding within VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。