测试发现顶尖视觉语言模型在基础视觉任务上表现堪忧。
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
- 对比模型各组件性能,识别出能力短板
- 部分基础任务准确率低于预期水平
- 适合研究视觉语言模型底层机制的学者
视觉语言模型(VLMs)已成为解决多种复杂计算机视觉问题的通用工具。尽管表现出色,但其在某些基础视觉理解任务上仍显不足。本文通过构建一系列测试,探究当前主流VLMs在基本视觉任务上的局限性,重点分析模型设计中的薄弱环节。不同于仅评估最终输出表现的现有基准,本研究还对比了直接基于视觉编码器、视觉-语言投影层及大语言模型解码器输出特征训练的探针模型性能。结果揭示了VLMs在能力、鲁棒性及视觉信息处理方式上的若干缺陷,为未来改进提供了关键洞见。
原文摘要 · Abstract (English)
Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, lacking some basic visual understanding skills. In this paper, we set out to understand the limitations of SoTA VLMs on fundamental visual tasks by constructing a series of tests that probe which components of design, specifically, may be lacking. Importantly, we go significantly beyond the current benchmarks, which simply measure the final performance of VLM response, by also comparing and contrasting it to the performance of probes trained directly on features obtained from the visual encoder, intermediate vision-language projection and LLM-decoder output. In doing so, we uncover shortcomings in VLMs and make a number of important observations about their capabilities, robustness and how they process visual information. We hope our insights will guide progress in further improving VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。