探究大模型如何理解声音与意义的关联,发现其具备类人音义联想能力。
Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound Symbolism
- 构建跨语言音义数据集LEX-ICON,涵盖8052个真实词和2930个伪词。
- 模型在25个语义维度上表现出与语言学研究一致的音义关联性。
- 通过注意力分析揭示模型关注具象化音素,适合认知科学与AI交叉研究者。
音义象征是语言中语音形式与意义之间非任意关联的概念。本文提出,这可作为探测多模态大语言模型(MLLMs)理解人类语言听觉信息的有效工具。研究考察了模型在文本(拼写与国际音标)和听觉输入下对语音象似性的表现,覆盖多达25个语义维度(如锐利与圆润)。通过测量音素层面的关注度分数,分析模型各层的信息处理过程。为此,本文构建了LEX-ICON,一个包含4种自然语言(英语、法语、日语、韩语)共8,052个真实词及2,930个系统构造伪词的大规模拟声词数据集,并在文本与音频模态上标注了语义特征。主要发现表明:(1)模型在多个语义维度上的语音直觉与现有语言学研究高度一致;(2)音义注意力模式凸显模型对具象化音素的关注。该成果首次实现了对MLLM音义解释能力的大规模定量分析,弥合人工智能与认知语言学之间的鸿沟。
原文摘要 · Abstract (English)
Sound symbolism is a linguistic concept that refers to non-arbitrary associations between phonetic forms and their meanings. We suggest that this can be a compelling probe into how Multimodal Large Language Models (MLLMs) interpret auditory information in human languages. We investigate MLLMs' performance on phonetic iconicity across textual (orthographic and IPA) and auditory forms of inputs with up to 25 semantic dimensions (e.g., sharp vs. round), observing models' layer-wise information processing by measuring phoneme-level attention fraction scores. To this end, we present LEX-ICON, an extensive mimetic word dataset consisting of 8,052 words from four natural languages (English, French, Japanese, and Korean) and 2,930 systematically constructed pseudo-words, annotated with semantic features applied across both text and audio modalities. Our key findings demonstrate (1) MLLMs' phonetic intuitions that align with existing linguistic research across multiple semantic dimensions and (2) phonosemantic attention patterns that highlight models' focus on iconic phonemes. These results bridge domains of artificial intelligence and cognitive linguistics, providing the first large-scale, quantitative analyses of phonetic iconicity in terms of MLLMs' interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。