arXiv:2510.14040cs.CL2025-10

跨6种语言量化语音与语义关联,发现新规律并验证旧假设。

Quantifying Phonosemantic Iconicity Distributionally in 6 Languages

  • 用分布方法分析6种语言中音素与语义空间的对齐度。
  • 发现多个未被记录的新语音语义关联模式,跨语言具一致性。
  • 验证部分既有假说,结果支持或存疑,适合语言学与计算研究者。

语言通常被认为是任意的,但语音与语义之间的系统性关联已在特定案例中被观察到。这种系统性关系在大规模、定量研究中能揭示多大程度?本文采用分布方法,在英语、西班牙语、印地语、芬兰语、土耳其语和泰米尔语这六种不同语言中,分析词素的语音相似性与语义相似性空间之间的对齐情况,使用一系列统计指标。研究发现了多个文献中未报道的可解释的语音语义对齐现象,并揭示了跨语言模式。同时,对五项已有假说进行检验,部分得到支持,部分结果呈现混合特征。

原文摘要 · Abstract (English)

Language is, as commonly theorized, largely arbitrary. Yet, systematic relationships between phonetics and semantics have been observed in many specific cases. To what degree could those systematic relationships manifest themselves in large scale, quantitative investigations--both in previously identified and unidentified phenomena? This work undertakes a distributional approach to quantifying phonosemantic iconicity at scale across 6 diverse languages (English, Spanish, Hindi, Finnish, Turkish, and Tamil). In each language, we analyze the alignment of morphemes' phonetic and semantic similarity spaces with a suite of statistical measures, and discover an array of interpretable phonosemantic alignments not previously identified in the literature, along with crosslinguistic patterns. We also analyze 5 previously hypothesized phonosemantic alignments, finding support for some such alignments and mixed results for others.

语音语义跨语言分布分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。