用数学函数解析词汇增长规律,打通统计模型与随机过程的联系。
Vocabulary Growth Fundamentals: Bernstein Functions and Hausdorff Sequences
- 通过伯恩斯坦函数和豪斯多夫序列建模词汇增长期望。
- 证明对数型偶发词率模型具有非负谱,属伯恩斯坦函数。
- 揭示理论局限性,拓展至平稳与威布尔更新过程场景。
我们综述了基于随机过程的词汇增长理论。特别地,利用伯恩斯坦函数和豪斯多夫序列来建模类型数量的期望值。这两类数学对象分别由导数或差分的交替符号定义,可分别对应连续时间泊松点过程和离散时间独立同分布过程。在已有词汇增长研究基础上,整合伯恩斯坦函数与豪斯多夫序列的更广义理论,并将其与近期发展的偶发词率模型相连接。具体而言,我们证明对数型偶发词率模型具有非负谱,因而属于伯恩斯坦函数,解决了此前提出的问题。同时,通过考虑平稳与威布尔更新过程下的推广,分析了伯恩斯坦-豪斯多夫理论在词汇增长中的局限性。
原文摘要 · Abstract (English)
We survey the theory of vocabulary growth founded in the setting of stochastic processes. In particular, we model the expected number of types through Bernstein functions and Hausdorff sequences. These classes of mathematical objects, defined by alternating signs of their derivatives or differences, can be related to continuous-time Poisson point processes and discrete-time IID processes, respectively. Building on previous accounts of the vocabulary growth, we integrate the broader theories of Bernstein functions and Hausdorff sequences and connect them with recently developed hapax rate models. In particular, we prove that the logistic hapax rate model has a non-negative spectrum and hence it defines a Bernstein function, thereby solving an earlier posed problem. We also analyze the limitations of the Bernstein--Hausdorff theory of the vocabulary growth by considering its generalizations under stationary and Weibull renewal processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。