修正赫德定律的二次项,更准描述词汇类型与数量关系
Quadratic Term Correction on Heaps' Law
- 在双对数坐标下引入二次项,拟合词频-词类曲线
- 线性系数略大于1,二次项系数约为-0.02
- 揭示真实数据中幂律关系的微小弯曲,适合语言统计研究者
赫德定律描述词类与词频间的幂律关系,在双对数坐标下呈直线。然而观察发现,该曲线仍存在轻微凹性,违背严格幂律。通过对二十部英文小说(部分为翻译作品)的数据分析,我们发现在双对数尺度下加入二次项可完美拟合词类-词频数据。回归分析显示,线性系数略大于1,二次系数约为-0.02。基于‘带替换抽球’模型,我们证明该凹性对应一个负的‘伪方差’。尽管当词频较大时伪方差计算可能数值不稳定,但该形式可在词频较小时提供曲率粗略估计。
原文摘要 · Abstract (English)
Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a straight line in log-log scale. However, it has been observed that even in log-log scale, the type-token curve is still slightly concave, invalidating the power-law relation. At the next-order approximation, we have shown, by twenty English novels or writings (some are translated from another language to English), that quadratic functions in log-log scale fit the type-token data perfectly. Regression analyses of log(type)-log(token) data with both a linear and quadratic term consistently lead to a linear coefficient of slightly larger than 1, and a quadratic coefficient around -0.02. Using the ``random drawing colored ball from the bag with replacement" model, we have shown that the curvature of the log-log scale is identical to a ``pseudo-variance" which is negative. Although a pseudo-variance calculation may encounter numeric instability when the number of tokens is large, due to the large values of pseudo-weights, this formalism provides a rough estimation of the curvature when the number of tokens is small.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。