大规模语音模型揭示深层语言关系,发现太平洋语言集群
Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster
- 扩展模型至4017种语言,发现非线性演化特征
- 4000语言模型首次识别出跨语系的太平洋宏观集群
- 模型捕捉到语音能量动态等深层声学共性,适合语言演化研究
自监督语音模型(S3Ms)提取的语言表征通常反映地理邻近或近期接触导致的表面类型相似性,可能忽略深层谱系信号。我们研究将基于S3M的语言识别系统从126种扩展至4,017种语言时的拓扑变化,发现存在非线性效应:在千级规模内谱系恢复保持平稳,但4000语言模型发生质变,同时解析出清晰的谱系关系与长期语言接触。最显著的是,一个稳定的太平洋宏集群浮现,将无亲缘关系的巴布亚、美拉尼西亚和澳大利亚语言归为一类。我们追溯其成因在于集中编码的声学特征,如全局能量动态。结果表明,大规模S3Ms可内化多层语言历史,为计算谱系学和语言接触研究提供新视角。
原文摘要 · Abstract (English)
Similarities between language representations derived from Self-Supervised Speech Models (S3Ms) have been observed to primarily reflect geographic proximity or surface typological similarities driven by recent expansion or contact, potentially missing deeper genealogical signals. We investigate how scaling an S3M-based language identification system from 126 to 4,017 languages reshapes this topology, and find a non-linear effect: phylogenetic recovery stays flat up to the 1K scale, but the 4K model undergoes a qualitative shift, resolving both clear lineages and long-term linguistic contact. Most strikingly, a robust Pacific macro-cluster emerges, grouping genealogically unrelated Papuan, Oceanic, and Australian languages, and we trace its driver to a concentrated encoding that captures shared acoustic signatures such as global energy dynamics. These results suggest that massive S3Ms internalize multiple layers of language history, offering a promising perspective for computational phylogenetics and the study of language contact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。