arXiv:2604.11565cs.CLcond-mat.stat-mech2026-04

用音位关联分析语言亲缘关系,推断印欧语系起源地。

Phonological distances for linguistic typology and the origin of Indo-European languages

  • 将音位序列建模为二阶马尔可夫链,捕捉语音系统统计相关性。
  • 67种语言的语音距离矩阵揭示语系归属与语言接触痕迹。
  • 语音距离与地理距离高度相关,支持印欧语起源于草原假说。

我们发现短程音位依赖关系能编码大规模语言亲缘模式,对量化类型学和演化语言学具有直接意义。具体而言,基于信息论框架,我们主张将音位序列建模为二阶马尔可夫链,可有效捕获语音系统的统计相关性。该方法使我们能够利用包含音位发音特征的距离度量,对来自多语言平行语料库的67种现代语言进行语音距离量化。所得语音距离矩阵恢复了主要语系结构,并揭示了语言接触导致的趋同迹象。值得注意的是,语音距离与地理距离存在显著相关性,从而可约束印欧语系可能的发源区域,结果与草原假说一致。

原文摘要 · Abstract (English)

We show that short-range phoneme dependencies encode large-scale patterns of linguistic relatedness, with direct implications for quantitative typology and evolutionary linguistics. Specifically, using an information-theoretic framework, we argue that phoneme sequences modeled as second-order Markov chains essentially capture the statistical correlations of a phonological system. This finding enables us to quantify distances among 67 modern languages from a multilingual parallel corpus employing a distance metric that incorporates articulatory features of phonemes. The resulting phonological distance matrix recovers major language families and reveals signatures of contact-induced convergence. Remarkably, we obtain a clear correlation with geographic distance, allowing us to constrain a plausible homeland region for the Indo-European family, consistent with the Steppe hypothesis.

语言演化语音分析印欧语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。