现有方法无法生成足够大的同源词数据集,制约了计算演化分析在语言学中的应用。
The Cognate Data Bottleneck in Language Phylogenetics
- 从BabelNet自动提取同源词数据
- 生成的数据构建的谱系树与标准树严重不符
- 揭示多语言资源难产高质量同源数据,适合历史语言学家参考
为充分发挥计算演化方法在同源词数据中的潜力,需依赖复杂模型与基于机器学习的技术。然而,这些方法要求的数据集远超当前人工收集的规模。据我们所知,尚无可行方法可自动生成更大规模的同源词数据集。本文通过从大型多语言百科词典BabelNet中自动提取数据加以验证,发现基于相应字符矩阵的演化推断结果与公认的标准谱系树存在显著不一致。同时讨论了从其他多语言资源中提取更合适字符矩阵的可能性极低。因此,需要大规模数据的演化分析方法目前无法应用于同源词数据。如何以及是否能将这些计算方法引入历史语言学,仍是一个开放问题。
原文摘要 · Abstract (English)
To fully exploit the potential of computational phylogenetic methods for cognate data one needs to leverage specific (complex) models an machine learning-based techniques. However, both approaches require datasets that are substantially larger than the manually collected cognate data currently available. To the best of our knowledge, there exists no feasible approach to automatically generate larger cognate datasets. We substantiate this claim by automatically extracting datasets from BabelNet, a large multilingual encyclopedic dictionary. We demonstrate that phylogenetic inferences on the respective character matrices yield trees that are largely inconsistent with the established gold standard ground truth trees. We also discuss why we consider it as being unlikely to be able to extract more suitable character matrices from other multilingual resources. Phylogenetic data analysis approaches that require larger datasets can therefore not be applied to cognate data. Thus, it remains an open question how, and if these computational approaches can be applied in historical linguistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。