arXiv:2509.19108cs.CL2025-09

检验语言中句子是否大多独特,发现文体影响显著。

Are most sentences unique? An empirical examination of Chomskyan claims

  • 用NLTK分析不同语料库,统计完全重复的句子
  • 多数语料中唯一句占多数,但文体差异大
  • 重复句在各类语料中均不可忽视,适合语言学研究者

语言学中普遍认为绝大多数语言表达都是独特的。例如,皮克林(1994)总结诺姆·乔姆斯基的观点称:‘一个人说出或理解的每句话几乎都是词的新组合,在宇宙史上首次出现。’ 随着大规模语料库的可用性提升,这一观点可被实证检验。本文利用NLTK Python库解析不同语体的语料库,统计其中完全相同的字符串匹配数量。结果显示,尽管在多数语料中完全独特的句子占多数,但这一比例受语体类型强烈影响,且重复句子在任何单一语料库中均非微不足道。研究揭示了语言使用中的独特性并非普遍恒定,而具有显著的语体依赖性。

原文摘要 · Abstract (English)

A repeated claim in linguistics is that the majority of linguistic utterances are unique. For example, Pinker (1994: 10), summarizing an argument by Noam Chomsky, states that "virtually every sentence that a person utters or understands is a brand-new combination of words, appearing for the first time in the history of the universe." With the increased availability of large corpora, this is a claim that can be empirically investigated. The current paper addresses the question by using the NLTK Python library to parse corpora of different genres, providing counts of exact string matches in each. Results show that while completely unique sentences are often the majority of corpora, this is highly constrained by genre, and that duplicate sentences are not an insignificant part of any individual corpus.

语言学语料分析重复句

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。