检验视觉语言模型是否像人一样理解形状与发音的关联。
Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
- 用提示词和Grad-CAM分析模型对形状-词汇匹配的反应。
- 两种模型对圆润形状偏好不一致,整体无稳定联想效应。
- 结果表明模型缺乏人类级跨模态认知,适合关注模型理解局限的研究者。
近期多模态模型的发展引发了关于视觉语言模型(VLMs)是否以反映人类认知的方式整合跨模态信息的疑问。一个经典测试案例是‘bouba-kiki’效应:人类会将伪词‘bouba’与圆润形状、‘kiki’与尖锐形状可靠关联。鉴于此前研究证据不一,本文对两种关键模型——基于ResNet和视觉变压器(ViT)的CLIP变体进行系统重评。采用两种贴近人类实验的方法:基于提示词的评估(以概率衡量模型偏好),以及使用Grad-CAM作为新方法解析形状-词汇匹配任务中的视觉注意力。结果显示,这两种模型并未表现出一致的bouba-kiki效应。尽管ResNet显示对圆润形状有偏好,但整体表现缺乏预期的关联模式。与先前人类数据直接对比表明,模型响应明显偏离人类具身认知中稳健的跨模态整合特征。这一结果为多模态模型是否真正理解跨模态概念的讨论提供了实证支持,揭示其内部表征与人类直觉之间的差距。
原文摘要 · Abstract (English)
Recent advances in multimodal models have raised questions about whether vision-and-language models (VLMs) integrate cross-modal information in ways that reflect human cognition. One well-studied test case in this domain is the bouba-kiki effect, where humans reliably associate pseudowords like `bouba' with round shapes and `kiki' with jagged ones. Given the mixed evidence found in prior studies for this effect in VLMs, we present a comprehensive re-evaluation focused on two variants of CLIP, ResNet and Vision Transformer (ViT), given their centrality in many state-of-the-art VLMs. We apply two complementary methods closely modelled after human experiments: a prompt-based evaluation that uses probabilities as a measure of model preference, and we use Grad-CAM as a novel approach to interpret visual attention in shape-word matching tasks. Our findings show that these model variants do not consistently exhibit the bouba-kiki effect. While ResNet shows a preference for round shapes, overall performance across both model variants lacks the expected associations. Moreover, direct comparison with prior human data on the same task shows that the models' responses fall markedly short of the robust, modality-integrated behaviour characteristic of human cognition. These results contribute to the ongoing debate about the extent to which VLMs truly understand cross-modal concepts, highlighting limitations in their internal representations and alignment with human intuitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。