揭示复杂系统中词频与词类数量的精确数学关系。
Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings
- 基于词频排名构建确定性模型,不依赖随机过程。
- 首次给出所有幂律指数α下的精确渐近公式。
- 适用于语言学、生态学等具幂律分布的系统研究。
复杂系统的增长常表现出幂律统计规律。真实有限系统中,类型计数S与类型排名r遵循反幂律关系S ∼ r⁻ᵃ,即齐普夫定律,在动物、词汇等场景中广泛存在。另一重要规律是希普斯定律:类型数随总词数亚线性增长。本文提出一种理想化增长模型,能确定性生成任意反幂律类型排名,并精确推导出类型-词数关系的渐近表达式。该方法统一处理所有α值,修正了α=1和α≫1时的特例问题,仅依赖排名形式,无需随机近似或抽样。结果表明,一般类型-词数关系可完全由齐普夫定律导出。
原文摘要 · Abstract (English)
The growth dynamics of complex systems often exhibit statistical regularities involving power-law relationships. For real finite complex systems formed by countable tokens (animals, words) as instances of distinct types (species, dictionary entries), an inverse power-law scaling $S \sim r^{-α}$ between type count $S$ and type rank $r$, widely known as Zipf's law, is widely observed to varying degrees of fidelity. A secondary, summary relationship is Heaps' law, which states that the number of types scales sublinearly with the total number of observed tokens present in a growing system. Here, we propose an idealized model of a growing system that (1) deterministically produces arbitrary inverse power-law count rankings for types, and (2) allows us to determine the exact asymptotics of the type-token relationship. Our argument improves upon and remedies earlier work. We obtain a unified asymptotic expression for all values of $α$, which corrects the special cases of $α= 1$ and $α\gg 1$. Our approach relies solely on the form of count rankings, avoids unnecessary approximations, and does not involve any stochastic mechanisms or sampling processes. We thereby demonstrate that a general type-token relationship arises solely as a consequence of Zipf's law.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。