arXiv:2411.10227cs.CLcs.IR2024-11被引 16

发现词频熵与词类比值的统一关系,揭示语言多样性的统计规律。

Entropy and type-token ratio in gigaword corpora

  • 通过六大数据集验证熵与词类比值的函数关系。
  • 在长文本下推导出符合齐普夫和希帕斯定律的解析表达式。
  • 适用于研究不同语言、文体的语言多样性,适合语言学与计算语言学研究者。

语言多样性可通过词类比值和词熵来衡量。本文研究了英语、西班牙语和土耳其语的六大数据集(涵盖书籍、新闻、推文),这些语料库具有不同的形态特征、语体和体裁,构成语言多样性的量化分析测试平台。研究发现,同一语料库和语言中,词熵与词类比值存在确定的函数关系,这是自然语言统计规律的结果。在文本长度趋近无穷时,基于齐普夫定律和希帕斯定律,推导出该关系的解析表达式,且与实证结果高度吻合。

原文摘要 · Abstract (English)

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six massive linguistic datasets in English, Spanish, and Turkish, consisting of books, news articles, and tweets. These gigaword corpora correspond to languages with distinct morphological features and differ in registers and genres, thus constituting a varied testbed for a quantitative approach to lexical diversity. We unveil an empirical functional relation between entropy and type-token ratio of texts of a given corpus and language, which is a consequence of the statistical laws observed in natural language. Further, in the limit of large text lengths we find an analytical expression for this relation relying on both Zipf and Heaps laws that agrees with our empirical findings.

语言多样性词熵类型比值统计语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。