arXiv:2609.05262cs.CL2026-09

无需标注即可快速构建全球语言演化树,仅用音标词表自监督学习。

Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference

论文配图:Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference
图 1 · 摘自论文原文
  • 通过双对比损失从原始音标词表中自监督学习词汇表征。
  • 在3,399种语言上构建进化树,计算仅需几分钟,性能媲美多个基线。
  • 表征同时支持语言演化与词义稳定性的下游分析,适合大规模语言研究者。

计算系统发育学已成为历史语言学的重要工具,但其在全球尺度的应用受限于两个因素:基于特征的方法需要耗时的手动同源判断标注,以及在大数据集上的推理计算成本高昂。本文提出一种完全自监督的对比学习框架,直接从原始IPA音标词表中学习词汇表征,无需同源标注、对齐或额外专家输入。模型采用双对比目标:词级损失将语音相似形式组织到统一空间,辅助语言级损失促使词汇空间反映语言的广义音系特性。基于生成的词表征,推导出成对语言距离并构建3,399种语言变体的全局系统发育树。该树在广义四元组距离(GQD)上与Glottolog参考树相比具有竞争力,且仅需标准笔记本GPU数分钟计算时间。此外,相同表征可捕捉历时概念稳定性:跨语言成对距离的方差生成稳定性排名,与既有排名显著相关。消融实验表明,语言级目标及使用音素特征向量均提升了树拓扑结构的GQD表现。该框架为大规模系统发育推断提供了高效、全自动替代方案,并提供统一表征以支持语言与概念层面的下游分析。

原文摘要 · Abstract (English)

Computational phylogenetics has become an essential tool in historical linguistics, yet its application at a global scale remains constrained by two factors: the labor-intensive manual annotation of cognacy judgments required for character-based methods and the substantial computational cost of inference on large datasets. This paper introduces a fully self-supervised contrastive learning framework that learns lexical representations directly from raw IPA-transcribed wordlists, without requiring cognacy annotations, alignments, or additional expert input. The model employs a dual contrastive objective: a word-level loss that organizes phonetically similar forms into a coherent space, and an auxiliary language-level loss that encourages the lexical space to reflect broader phonological properties of languages. From the resulting word representations, pairwise language distances are derived and used to infer a global phylogenetic tree of 3,399 language varieties. The inferred tree achieves a generalized quartet distance (GQD) to the Glottolog reference tree competitive with multiple baselines, while requiring only minutes of computation on a standard notebook GPU. Furthermore, the same representations capture diachronic concept stability: variance in pairwise distances across languages yields stability rankings that correlate significantly with established rankings. Ablation studies confirm that both the language-level objective and the use of phonetic feature vectors improved the inferred trees topology with regards to GQD. The framework thus provides a computationally efficient and fully automatic alternative for large-scale phylogenetic inference and offers a unified representation supporting downstream analyses at both the language and concept level.

系统发育自监督学习语言演化词表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。