用神经网络自动判断语言亲缘关系,提升历史语言学研究效率。
From Isolates to Families: Using Neural Networks for Automated Language Affiliation
- 基于1000+种语言的词汇与语法数据训练模型,实现自动分类。
- 仅用词汇数据的模型表现优于纯语法模型,二者结合效果最佳。
- 可帮助识别孤立语言的潜在亲属,探索未知语言归属。
在历史语言学中,语言归属到共同语系的传统方法依赖繁琐的人工比对。大规模标准化多语言词表和语法结构数据可能推动自动化归属流程的发展。本文提出神经网络模型,利用全球超过1000种已知语系归属语言的词汇与语法数据,对单个语言进行语系分类。结果表明,仅使用词汇数据的模型表现优于仅依赖语法数据的模型,而两者结合可进一步提升性能。额外实验显示,该模型能识别大范围语群间的关联,可用于探究孤立语言的潜在亲属,并为尚未归属的语言提供初步归属线索。结论表明,基于词汇与语法数据训练的自动化语言归属模型,为比较语言学家评估深层未知语言关系提供了有力工具。
原文摘要 · Abstract (English)
In historical linguistics, the affiliation of languages to a common language family is traditionally carried out using a complex workflow that relies on manually comparing individual languages. Large-scale standardized collections of multilingual wordlists and grammatical language structures might help to improve this and open new avenues for developing automated language affiliation workflows. Here, we present neural network models that use lexical and grammatical data from a worldwide sample of more than 1,000 languages with known affiliations to classify individual languages into families. In line with the traditional assumption of most linguists, our results show that models trained on lexical data alone outperform models solely based on grammatical data, whereas combining both types of data yields even better performance. In additional experiments, we show how our models can identify long-ranging relations between entire subgroups, how they can be employed to investigate potential relatives of linguistic isolates, and how they can help us to obtain first hints on the affiliation of so far unaffiliated languages. We conclude that models for automated language affiliation trained on lexical and grammatical data provide comparative linguists with a valuable tool for evaluating hypotheses about deep and unknown language relations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。