用视觉语言模型测试声音象征性,发现模型能听懂声音与概念的非任意联系。
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
- 用多模态模型模拟人类听觉感知,通过图像和文字推断声音象征关系。
- 模型对大小象征的判断准确率高于形状象征,且大模型表现更优。
- 适合对认知科学、多模态模型能力感兴趣的读者参考。
大型语言模型(LLMs)和视觉语言模型(VLMs)在模拟人类参与心理语言学实验方面展现出潜力。然而,仅具备视觉与文本模态的模型,能否仅通过书写形式和图像进行抽象推理来隐式理解基于声音的现象,仍缺乏研究。为此,我们分析了VLMs和LLMs在声音象征性任务中的表现,包括复现经典的Kiki-Bouba和Mil-Mal形状与大小象征任务,并对比人类与模型对语言象似性的判断。结果显示,VLMs与人类标签存在不同程度的一致性,且其完成虚拟实验所需的任务信息量高于人类;更大规模的模型在识别语言象似性方面表现更佳。此外,大小象征比形状象征更容易被模型识别。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated aptitude as potential substitutes for human participants in experiments testing psycholinguistic phenomena. However, an understudied question is to what extent models that only have access to vision and text modalities are able to implicitly understand sound-based phenomena via abstract reasoning from orthography and imagery alone. To investigate this, we analyse the ability of VLMs and LLMs to demonstrate sound symbolism (i.e., to recognise a non-arbitrary link between sounds and concepts) as well as their ability to "hear" via the interplay of the language and vision modules of open and closed-source multimodal models. We perform multiple experiments, including replicating the classic Kiki-Bouba and Mil-Mal shape and magnitude symbolism tasks, and comparing human judgements of linguistic iconicity with that of LLMs. Our results show that VLMs demonstrate varying levels of agreement with human labels, and more task information may be required for VLMs versus their human counterparts for in silico experimentation. We additionally see through higher maximum agreement levels that Magnitude Symbolism is an easier pattern for VLMs to identify than Shape Symbolism, and that an understanding of linguistic iconicity is highly dependent on model size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。