arXiv:2504.10786cs.CVcs.AI2025-04被引 5

VLMs在基础视觉任务上表现堪忧,像人类患者一样存在明显缺陷。

Visual Language Models show widespread visual deficits on neuropsychological tests

  • 用神经心理学测试工具系统评估3个主流VLM的视觉能力
  • 51项测试中多数模型在方向、位置等低层视觉任务上显著落后于人类
  • 虽能识别复杂物体,但缺乏人类无需训练的基础视觉概念

视觉语言模型(VLMs)在视觉推理任务中表现出色,能解决需高级图像理解的大学水平问题。然而,近期报告指出这些模型在方向、位置、连续性、遮挡等基本视觉概念推理上存在困难,暗示其与人类视觉之间可能存在鸿沟。本文采用神经心理学工具包,系统评估三个先进VLM在多个视觉领域的表现。基于6个临床与实验量表中的51项测试,我们对比了领先VLM与健康成年人的正常表现。尽管模型在简单物体识别任务中表现优异,但在低至中等层次的视觉能力上普遍存在缺陷,若用于人类,这些缺陷会被视为临床上显著的问题。通过标准化测试量表所描绘的特征性缺陷表明,一个人工智能系统可在无需显式训练的情况下实现复杂物体识别,却仍缺乏人类所需的基础视觉概念。

原文摘要 · Abstract (English)

Visual Language Models (VLMs) show remarkable performance in visual reasoning tasks, successfully tackling college-level challenges that require high-level understanding of images. However, some recent reports of VLMs struggling to reason about elemental visual concepts like orientation, position, continuity, and occlusion suggest a potential gulf between human and VLM vision. Here we use the toolkit of neuropsychology to systematically assess the capabilities of three state-of-the-art VLMs across visual domains. Using 51 tests drawn from six clinical and experimental batteries, we characterise the visual abilities of leading VLMs relative to normative performance in healthy adults. While the models excel in straightforward object recognition tasks, we find widespread deficits in low- and mid-level visual abilities that would be considered clinically significant in humans. These selective deficits, profiled through validated test batteries, suggest that an artificial system can achieve complex object recognition without developing foundational visual concepts that in humans require no explicit training.

视觉语言模型神经心理学视觉缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。