从词频分布规律推导出大模型训练规律,揭示语言统计与模型性能的深层联系。
From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- 由词频幂律推出词汇量增长规律,再推导出信息熵变化趋势
- 在假设条件下,证明神经网络缩放定律可由词频规律推导得出
- 适用于理解大模型性能与语言统计特性的理论研究者
我们系统考察了神经网络缩放定律与齐普夫定律之间的演绎关系——这两个定律分别在机器学习和定量语言学中被广泛讨论。神经缩放定律描述了基础模型(如大语言模型)的交叉熵随训练语料量、参数量和计算量的变化规律;而齐普夫定律指出词汇分布具有幂律尾部特征。尽管已有研究在特定场景下提出类似观点,本文在一系列广义假设下,首次明确推导出神经缩放定律是齐普夫定律的直接结果。推导路径为:由齐普夫定律推导海普斯定律(词汇量增长),再由海普斯定律推导希尔伯格假说(熵的标度行为),最终得到神经缩放定律。我们通过一个满足四种统计规律的圣塔菲过程简化示例验证了该推导链条的逻辑一致性。
原文摘要 · Abstract (English)
We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate of a foundation model -- such as a large language model -- changes with respect to the amount of training tokens, parameters, and compute. By contrast, Zipf's law posits that the distribution of tokens exhibits a power law tail. Whereas similar claims have been made in more specific settings, we show that the neural scaling law is a consequence of Zipf's law under certain broad assumptions that we reveal systematically. The derivation steps are as follows: We derive Heaps' law on the vocabulary growth from Zipf's law, Hilberg's hypothesis on the entropy scaling from Heaps' law, and the neural scaling from Hilberg's hypothesis. We illustrate these inference steps by a toy example of the Santa Fe process that satisfies all four statistical laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。