通过眼科检查法分析视觉语言模型的感知能力,发现其对绿色普遍不敏感。
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
- 设计LIVE检查流程,系统评估模型对颜色、形状和语义的感知能力。
- 发现所有模型对绿色不敏感,且不同模型在形状和语义识别上表现差异显著。
- 结果可指导模型设计与视觉输入预处理,提升实际应用性能。
视觉语言模型(VLMs)在多个基准测试中展现出令人瞩目的推理能力,但对其视觉感知的理解仍有限。本文提出一种‘眼科检查’流程,用于探究VLM如何感知图像,重点关注从基础的颜色、形状到语义层面的视觉识别能力。为此,我们构建了一个名为LENS的数据集,引导VLM完成检查并评估其准备状态。当模型就绪后,开展系统性检查。通过该流程,我们量化并可视化了VLM对颜色、形状及语义匹配的敏感度。研究发现,不同模型对各类颜色的敏感度各异,但均表现出对绿色的持续不敏感;同时,尽管使用相同的固定视觉编码器,大语言模型(LLM)容量差异也导致形状敏感性和语义识别能力的不同。这些分析与发现为提升VLM设计及视觉输入预处理提供了潜在方向。
原文摘要 · Abstract (English)
Vision language models (VLMs) have shown promising reasoning capabilities across various benchmarks; however, our understanding of their visual perception remains limited. In this work, we propose an eye examination process to investigate how a VLM perceives images, specifically focusing on key elements of visual recognition, from primitive color and shape to semantic levels. To this end, we introduce a dataset named LENS to guide a VLM to follow the examination and check its readiness. Once the model is ready, we conduct the examination. Through this examination, we quantify and visualize VLMs' sensitivities to color and shape, and semantic matching. Our findings reveal that VLMs have varying sensitivity to different colors while consistently showing insensitivity to green across different VLMs. Also, we found different shape sensitivity and semantic recognition depending on LLM's capacity despite using the same fixed visual encoder. Our analyses and findings have potential to inspire the design of VLMs and the pre-processing of visual input to VLMs for improving application performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。