arXiv:2512.12245cs.CLcs.AI2025-12

跨27种语言验证声音与意义的非任意关联,发现音素可预测大小语义

Adversarially Probing Cross-Family Sound Symbolism in 27 Languages

  • 用音段特征构建可解释分类器,分析声音与意义关系
  • 跨语言音素能显著预测大小语义,即使语言无亲缘关系
  • 设计对抗清洗模型,证实跨语系声音象征性普遍存在

声音象征性(即语音与意义间的非任意映射)虽在如Bouba Kiki等实验中被广泛提及,但缺乏大规模实证。本文首次开展跨语言、大样本的声音象征性研究,聚焦尺寸语义领域。构建了包含810个形容词的语料库(27种语言,每种30个词),所有词汇均进行音位转写,并经母语者音频验证。采用基于音段袋模型的可解释分类器,发现语音形式在跨语言情境下仍能显著预测尺寸语义,且元音与辅音均有贡献。为检验跨语系普遍性,训练对抗性清洗模型,在抑制语言身份信息的同时保留尺寸信号(支持按语族粒度保留)。结果显示,语言识别准确率低于随机水平,而尺寸预测仍显著高于随机,证明存在跨语族声音象征偏倚。数据、代码及诊断工具已开源,支持未来大规模拟态性研究。

原文摘要 · Abstract (English)

The phenomenon of sound symbolism, the non-arbitrary mapping between word sounds and meanings, has long been demonstrated through anecdotal experiments like Bouba Kiki, but rarely tested at scale. We present the first computational cross-linguistic analysis of sound symbolism in the semantic domain of size. We compile a typologically broad dataset of 810 adjectives (27 languages, 30 words each), each phonemically transcribed and validated with native-speaker audio. Using interpretable classifiers over bag-of-segment features, we find that phonological form predicts size semantics above chance even across unrelated languages, with both vowels and consonants contributing. To probe universality beyond genealogy, we train an adversarial scrubber that suppresses language identity while preserving size signal (also at family granularity). Language prediction averaged across languages and settings falls below chance while size prediction remains significantly above chance, indicating cross-family sound-symbolic bias. We release data, code, and diagnostic tools for future large-scale studies of iconicity.

声音象征性跨语言语音学可解释模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。