视觉语言模型倾向于人类基本层级分类,反映其认知能力。
Basic Category Usage in Vision Language Models
- 模型在图像分类中偏好人类常用的基本层级类别
- 对生物与非生物类别的区分符合人类行为模式
- 适合研究大模型认知机制或人机共情设计的读者
心理学长期认为人类在标记视觉刺激时存在一个基本层级的分类方式,该概念由Rosch于1976年提出。此层级分类使用频率最高、信息密度更高,并有助于视觉语言任务中的启动效应。本文研究了两个新发布的开源视觉语言模型(Llama 3.2 Vision Instruct 11B 和 Molmo 7B-D)中的基本层级分类现象。结果表明,这两个模型均表现出与人类行为一致的基本层级偏好。此外,模型对生物与非生物类别的处理也符合人类的细微行为差异,如专家基本层级偏移现象,说明这些模型在训练数据中习得了复杂的认知分类能力。我们还发现,专家提示方法的准确率低于非专家提示方法,与当前普遍观点相悖。
原文摘要 · Abstract (English)
The field of psychology has long recognized a basic level of categorization that humans use when labeling visual stimuli, a term coined by Rosch in 1976. This level of categorization has been found to be used most frequently, to have higher information density, and to aid in visual language tasks with priming in humans. Here, we investigate basic-level categorization in two recently released, open-source vision-language models (VLMs). This paper demonstrates that Llama 3.2 Vision Instruct (11B) and Molmo 7B-D both prefer basic-level categorization consistent with human behavior. Moreover, the models' preferences are consistent with nuanced human behaviors like the biological versus non-biological basic level effects and the well-established expert basic level shift, further suggesting that VLMs acquire complex cognitive categorization behaviors from the human data on which they are trained. We also find our expert prompting methods demonstrate lower accuracy then our non-expert prompting methods, contradicting popular thought regarding the use of expertise prompting methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。