用几何机制解释语言中齐普夫定律的成因,无需依赖语义。
Zipf Distributions from Two-Stage Symbolic Processes: Stability Under Stochastic Lexical Filtering
- 基于符号组合的双阶段模型生成词长几何分布
- 词频排名呈幂律分布,由字母表大小和空符概率决定
- 模拟结果匹配英俄语及混合文本数据,适合复杂系统研究者
语言中的齐普夫定律尚无定论,跨学科争论不断。本研究通过无语言元素的几何机制解释齐普夫型行为。全组合词模型(FCWM)从有限字母表生成单词,导致词长服从几何分布。相互作用的指数力产生幂律型频率排名曲线,其形式由字母表大小和空白符号概率决定。模拟结果支持理论预测,与英语、俄语及混合题材文本数据吻合。该符号模型表明,齐普夫型规律源于几何约束,而非交际效率。
原文摘要 · Abstract (English)
Zipf's law in language lacks a definitive origin, debated across fields. This study explains Zipf-like behavior using geometric mechanisms without linguistic elements. The Full Combinatorial Word Model (FCWM) forms words from a finite alphabet, generating a geometric distribution of word lengths. Interacting exponential forces yield a power-law rank-frequency curve, determined by alphabet size and blank symbol probability. Simulations support predictions, matching English, Russian, and mixed-genre data. The symbolic model suggests Zipf-type laws arise from geometric constraints, not communicative efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。