arXiv:2502.06342cs.CLphysics.soc-ph2025-02被引 2

24种名词短语顺序分布呈指数型,挑战了幂律分布必然性。

The exponential distribution of the order of demonstrative, numeral, adjective and noun

  • 用指数分布模型拟合24种词序频率,优于幂律分布。
  • 所有24种顺序均有非零概率的指数模型拟合效果更优。
  • 支持词序无硬性约束,未出现顺序源于数据不足。

过去二十年间,由指示词、数词、形容词和名词构成的名词短语中,偏好词序的频率备受关注。本文研究24种可能词序的实际分布,探讨其是否符合指数分布或幂律分布。结果表明,指数分布是更优模型。这一发现及其他类似指数分布现象,挑战了幂律分布(如词汇频率的齐普夫定律)不可避免的观点。我们进一步比较两种指数模型:一种为24个词序均具有非零概率(在24位截断的几何分布);另一种为可变非零概率数量(右截断几何分布)。当强调一致性与泛化能力时,前者获得更高支持。研究强烈表明词序变化无硬性约束,未观测到的词序仅因采样不足所致,与Cysouw的观点一致。

原文摘要 · Abstract (English)

The frequency of the preferred order for a noun phrase formed by demonstrative, numeral, adjective and noun has received significant attention over the last two decades. We investigate the actual distribution of the 24 possible orders. There is no consensus on whether it is well-fitted by an exponential or a power law distribution. We find that an exponential distribution is a much better model. This finding and other circumstances where an exponential-like distribution is found challenge the view that power-law distributions, e.g., Zipf's law for word frequencies, are inevitable. We also investigate which of two exponential distributions gives a better fit: an exponential model where the 24 orders have non-zero probability (a geometric distribution truncated at rank 24) or an exponential model where the number of orders that can have non-zero probability is variable (a right-truncated geometric distribution). When consistency and generalizability are prioritized, we find higher support for the exponential model where all 24 orders have non-zero probability. These findings strongly suggest that there is no hard constraint on word order variation and then unattested orders merely result from undersampling, consistently with Cysouw's view.

语言学分布模型词序统计规律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。