随机文本也能生成类似语言的词频分布,揭示了大模型统计规律的底层机制。
Random Text, Zipf's Law, Critical Length,and Implications for Large Language Models
- 用独立符号生成文本,通过组合数学推导出词长和词汇量规律。
- 发现词长服从几何分布,存在关键长度使词汇从高频变为唯一出现。
- 无需语义或语法,仅靠随机性和分词规则即可产生类齐夫定律的词频结构。
我们研究一种完全非语言的简单文本模型:从有限字母表(含空格)中独立抽取符号序列。单词定义为连续非空格符号的最大块。在此符号级框架下,不依赖形态、句法或语义,我们得出若干结构结果。首先,词长服从几何分布,仅由空格出现概率决定。其次,给定长度的词数量及其不同词的数量可基于优惠券收集论推导出闭式表达,由此得到临界词长k*,即在此长度之上,平均每个词只出现一次。第三,结合长度为k的字符串数量的指数增长与每串概率的指数衰减,导出类齐夫定律的词频分布p(r) ∝ r^{-α},其中指数α由字母表大小和空格概率明确决定。本工作在数学上统一推导了词长、词汇增长、临界长度与秩频结构;在概念上提出该模型是自然语言词统计和大模型词元统计的结构性零模型。结果表明,齐夫型模式可纯粹由组合学与分词规则产生,无需优化或语言组织,有助于厘清哪些现象需超越随机结构的深层解释。
原文摘要 · Abstract (English)
We study a deliberately simple, fully non-linguistic model of text: a sequence of independent draws from a finite alphabet of letters plus a single space symbol. A word is defined as a maximal block of non-space symbols. Within this symbol-level framework, which assumes no morphology, syntax, or semantics, we derive several structural results. First, word lengths follow a geometric distribution governed solely by the probability of the space symbol. Second, the expected number of words of a given length, and the expected number of distinct words of that length, admit closed-form expressions based on a coupon-collector argument. This yields a critical word length k* at which word types transition from appearing many times on average to appearing at most once. Third, combining the exponential growth of the number of possible strings of length k with the exponential decay of the probability of each string, we obtain a Zipf-type rank-frequency law p(r) proportional to r^{-alpha}, with an exponent determined explicitly by the alphabet size and the space probability. Our contribution is twofold. Mathematically, we give a unified derivation linking word lengths, vocabulary growth, critical length, and rank-frequency structure in a single explicit model. Conceptually, we argue that this provides a structurally grounded null model for both natural-language word statistics and token statistics in large language models. The results show that Zipf-like patterns can arise purely from combinatorics and segmentation, without optimization principles or linguistic organization, and help clarify which phenomena require deeper explanation beyond random-text structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。