arXiv:2510.22220cs.CLq-bio.PE2025-10被引 1

从概率视角重新审视词汇演化,提升语言年代推断精度。

Evolution of the lexicon: a probabilistic point of view

  • 引入词汇渐进变化的随机过程,弥补传统方法缺陷
  • 理论证明仅靠替换率无法完全准确推断语言年代
  • 结合两种随机过程可显著提高语言分化时间估计精度

基于词汇替换的斯瓦迪士语言年代推断方法依赖于词义更替的随机过程。然而,该方法常因横向传播、替换率时空变异、同源关系误判、同义词存在等干扰因素而失效。更重要的是,即使假设所有基本前提成立,单纯数学分析表明语言年代推断仍存在固有的概率性误差。本文详细分析了这些纯概率性限制,并指出语言词汇演化还受另一种随机过程驱动:词汇的渐进式修改。从概率视角出发,我们证明这一过程对词汇重塑具有重要贡献,且同时考虑两类随机过程可显著提升语言分化时间估计的精度。

原文摘要 · Abstract (English)

The Swadesh approach for determining the temporal separation between two languages relies on the stochastic process of words replacement (when a complete new word emerges to represent a given concept). It is well known that the basic assumptions of the Swadesh approach are often unrealistic due to various contamination phenomena and misjudgments (horizontal transfers, variations over time and space of the replacement rate, incorrect assessments of cognacy relationships, presence of synonyms, and so on). All of this means that the results cannot be completely correct. More importantly, even in the unrealistic case that all basic assumptions are satisfied, simple mathematics places limits on the accuracy of estimating the temporal separation between two languages. These limits, which are purely probabilistic in nature and which are often neglected in lexicostatistical studies, are analyzed in detail in this article. Furthermore, in this work we highlight that the evolution of a language's lexicon is also driven by another stochastic process: gradual lexical modification of words. We show that this process equally also represents a major contribution to the reshaping of the vocabulary of languages over the centuries and we also show, from a purely probabilistic perspective, that taking into account this second random process significantly increases the precision in determining the temporal separation between two languages.

词汇演化概率模型语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。