用词素组合机制解释词汇频率分布,无需意义或优化假设。
The Morphemic Origin of Zipf's Law: A Factorized Combinatorial Framework
- 基于词素位置槽的随机激活生成单词
- 模拟出长度分布与频率秩曲线均接近真实语言
- 揭示形态结构本身可导致齐普夫定律现象
我们提出一个基于词素结构的单词生成模型。该模型仅通过前缀、词根、后缀和变词语素的组合规则即可解释两个核心语言现象:词汇长度分布特征与齐普夫式频率秩曲线。与传统依赖随机文本或通信效率的观点不同,本模型通过为每个词素位置分配激活概率,并从对应词素库中选择一个,来生成单词。词素被视为稳定出现且具有固定位置的构建单元。该机制产生的词汇长度分布呈集中中间区与稀疏长尾特征,与真实语言高度一致。使用合成词素库进行的模拟生成了指数约为1.1-1.4的齐普夫式频率曲线,与英语、俄语及罗曼语族相似。关键结论是:齐普夫行为可在无意义、无通信压力、无优化原则的情况下,仅由形态结构与位置槽的随机激活自然产生。
原文摘要 · Abstract (English)
We present a simple structure based model of how words are formed from morphemes. The model explains two major empirical facts: the typical distribution of word lengths and the appearance of Zipf like rank frequency curves. In contrast to classical explanations based on random text or communication efficiency, our approach uses only the combinatorial organization of prefixes, roots, suffixes and inflections. In this Morphemic Combinatorial Word Model, a word is created by activating several positional slots. Each slot turns on with a certain probability and selects one morpheme from its inventory. Morphemes are treated as stable building blocks that regularly appear in word formation and have characteristic positions. This mechanism produces realistic word length patterns with a concentrated middle zone and a thin long tail, closely matching real languages. Simulations with synthetic morpheme inventories also generate rank frequency curves with Zipf like exponents around 1.1-1.4, similar to English, Russian and Romance languages. The key result is that Zipf like behavior can emerge without meaning, communication pressure or optimization principles. The internal structure of morphology alone, combined with probabilistic activation of slots, is sufficient to create the robust statistical patterns observed across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。