用信息论解释全球语言中音素频率分布规律。
The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models
- 宏观上用狄利克雷分布解释音素频次排名,库存越大熵越低。
- 微观上最大熵模型精准预测各语言音素概率,结合发音、音系等约束。
- 首次统一揭示音素分布的宏观与微观机制,适合语言学与计算研究者。
我们证明了全球语言中音素频率分布可在宏观和微观层面得到解释。宏观上,音素频次排名分布紧密符合对称狄利克雷分布的顺序统计特性,其单一集中参数随音素库大小系统性变化,揭示出一种稳健的补偿效应:音素库越大,相对熵越低。微观上,一个引入发音、音系及词汇结构约束的最大熵模型能准确预测特定语言的音素概率。两者共同为音素频率结构提供了统一的信息论解释。
原文摘要 · Abstract (English)
We demonstrate that the frequency distribution of phonemes across languages can be explained at both macroscopic and microscopic levels. Macroscopically, phoneme rank-frequency distributions closely follow the order statistics of a symmetric Dirichlet distribution whose single concentration parameter scales systematically with phonemic inventory size, revealing a robust compensation effect whereby larger inventories exhibit lower relative entropy. Microscopically, a Maximum Entropy model incorporating constraints from articulatory, phonotactic, and lexical structure accurately predicts language-specific phoneme probabilities. Together, these findings provide a unified information-theoretic account of phoneme frequency structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。