用自动比对方法构建语言演化树,突破人工标注瓶颈。
Beyond cognacy
- 用马尔可夫模型自动对齐词汇数据,提取演化信号
- 新方法生成的树更贴近语言分类,预测语言特征更准
- 适合大规模语言演化研究,无需依赖专家标注
计算谱系学已成为历史语言学的成熟工具,许多语系通过基于似然的推断方法分析。然而,传统方法依赖人工标注的同源词集,存在稀疏、耗时且仅限于个别语系的问题。本文对比了传统方法与两种全自动方法:一种基于单字/概念特征的自动同源词聚类;另一种基于配对隐马尔可夫模型的多序列比对(MSA)。两者在Glottolog专家分类和Grambank类型学数据上进行评估。结果表明,基于MSA的方法生成的系统发育树与语言分类更一致,对类型学变异的预测能力更强,且提供的演化信号更清晰,表明其是传统同源词方法的一种有前景的可扩展替代方案,为突破人工标注瓶颈、实现全球尺度语言谱系研究开辟新路径。
原文摘要 · Abstract (English)
Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-based inference. However, standard approaches rely on expert-annotated cognate sets, which are sparse, labor-intensive to produce, and limited to individual language families. This paper explores alternatives by comparing the established method to two fully automated methods that extract phylogenetic signal directly from lexical data. One uses automatic cognate clustering with unigram/concept features; the other applies multiple sequence alignment (MSA) derived from a pair-hidden Markov model. Both are evaluated against expert classifications from Glottolog and typological data from Grambank. Also, the intrinsic strengths of the phylogenetic signal in the characters are compared. Results show that MSA-based inference yields trees more consistent with linguistic classifications, better predicts typological variation, and provides a clearer phylogenetic signal, suggesting it as a promising, scalable alternative to traditional cognate-based methods. This opens new avenues for global-scale language phylogenies beyond expert annotation bottlenecks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。