用简单随机过程解释语言中整体与部分长度关系的规律。
Simple stochastic processes behind Menzerath's Law
- 假设词长变化在音节和音素层面独立且乘法性,导出双变量对数正态分布。
- 引入高斯耦合器建模联合分布,比经典模型更贴合真实数据。
- 适合语言统计学、计算语言学研究者参考,理解语言结构生成机制。
本文重新审视了著名的梅纳扎特定律(即梅纳扎特-阿尔特曼定律),该定律描述语言单位整体长度与其组成部分平均长度之间的关系。最新研究表明,简单的随机过程也能表现出梅纳扎特特性,但现有模型难以准确反映现实数据。若假设一个词在音节数和音素数上均可发生长度变化,且两者相关性不完全,变化具有乘法性质,则可推导出双变量对数正态分布。本文证明,仅从这一简单原则出发,即可导出经典的阿尔特曼模型。若将联合分布与边缘分布分别独立建模,采用高斯耦合器可构建更精确的模型。模型经实证数据验证,并讨论了其他替代方法。
原文摘要 · Abstract (English)
This paper revisits Menzerath's Law, also known as the Menzerath-Altmann Law, which models a relationship between the length of a linguistic construct and the average length of its constituents. Recent findings indicate that simple stochastic processes can display Menzerathian behaviour, though existing models fail to accurately reflect real-world data. If we adopt the basic principle that a word can change its length in both syllables and phonemes, where the correlation between these variables is not perfect and these changes are of a multiplicative nature, we get bivariate log-normal distribution. The present paper shows, that from this very simple principle, we obtain the classic Altmann model of the Menzerath-Altmann Law. If we model the joint distribution separately and independently from the marginal distributions, we can obtain an even more accurate model by using a Gaussian copula. The models are confronted with empirical data, and alternative approaches are discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。