arXiv:2602.02221cs.CL2026-02中稿 · the L'Change works…

提出新方法定量识别词源集中的异常词形,准确率达85%。

Using Correspondence Patterns to Identify Irregular Words in Cognate sets Through Leave-One-Out Validation

  • 用留一验证法评估音变模式的重复频率来量化规则性。
  • 在真实数据集上识别异常词形的准确率达85%。
  • 适合语言学家和计算语言学研究者用于提升词源数据质量。

音变规律是历史语言比较的核心证据。尽管规律性常被视为直观判断,而非量化评估,且异常现象比新格拉马尔理论预期更普遍。借助计算语言学进展与标准化词汇数据的普及,我们可实现更精确的评估。本文提出一种新的规则性度量——对应模式的平衡平均重复率,并基于此开发了一种新计算方法,用于识别不符合音变规律的词源集。通过模拟数据和真实数据的双实验验证,采用留一验证法,在替换一个词形为异常形式后,检测该方法能否准确识别出异常词。在真实数据集上,整体识别准确率达85%。同时展示了使用大数据子样本的优势,以及异常程度对结果的影响。我们认为该规则性度量及异常词识别方法,有望显著提升现有与未来计算机辅助语言比较数据集的质量。

原文摘要 · Abstract (English)

Regular sound correspondences constitute the principal evidence in historical language comparison. Despite the heuristic focus on regularity, it is often more an intuitive judgement than a quantified evaluation, and irregularity is more common than expected from the Neogrammarian model. Given the recent progress of computational methods in historical linguistics and the increased availability of standardized lexical data, we are now able to improve our workflows and provide such a quantitative evaluation. Here, we present the balanced average recurrence of correspondence patterns as a new measure of regularity. We also present a new computational method that uses this measure to identify cognate sets that lack regularity with respect to their correspondence patterns. We validate the method through two experiments, using simulated and real data. In the experiments, we employ leave-one-out validation to measure the regularity of cognate sets in which one word form has been replaced by an irregular one, checking how well our method identifies the forms causing the irregularity. Our method achieves an overall accuracy of 85\% with the datasets based on real data. We also show the benefits of working with subsamples of large datasets and how increasing irregularity in the data influences our results. Reflecting on the broader potential of our new regularity measure and the irregular cognate identification method based on it, we conclude that they could play an important role in improving the quality of existing and future datasets in computer-assisted language comparison.

历史语言学词源分析计算语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。