arXiv:2512.11880math.HOcs.CL2025-12

用信息论估算:有知识的猴子比随机打字猴快10^60倍产出莎士比亚语句

Shakespeare, Entropy and Educated Monkeys

  • 用统计典型文本约束随机打字,大幅减少无效尝试
  • 莎士比亚一句台词仅需7.3万年(原需2.7×10^63年)
  • 揭示自然语言生成的本质是压缩而非随机搜索

众所周知,一只永久随机打字的猴子最终会写出莎士比亚全集,但所需时间远超宇宙寿命。本文指出:若让一只‘受过教育’的猴子仍随机打字,但只允许生成‘统计上典型’的文本,则可显著缩短生成特定文本的时间。信息论提供了一种简单方法估算该时间。例如,莎士比亚《温莎的风流妇人》中的一句‘Better three hours too soon than a minute too late’,普通猴子需2.7×10^63年,而受教育猴子只需7.3万年。尽管如此,要生成整部《哈姆雷特》,仍需10^42,277年——依然远超任何实际意义的时间尺度。

原文摘要 · Abstract (English)

It has often been said, correctly, that a monkey forever randomly typing on a keyboard would eventually produce the complete works of William Shakespeare. Almost just as often it has been pointed out that this "eventually" is well beyond any conceivably relevant time frame. We point out that an educated monkey that still types at random but is constrained to only write "statistically typical" text, would produce any given passage in a much shorter time. Information theory gives a very simple way to estimate that time. For example, Shakespeare's phrase, Better three hours too soon than a minute too late, from The Merry Wives of Windsor, would take the educated monkey only 73 thousand years to produce, compared to the beyond-astronomical $2.7 \times 10^{63}$ years for the randomly typing one. Despite the obvious improvement, it would still take the educated monkey an unimaginably long $10^{42,277}$ years to produce all of Hamlet.

信息论自然语言生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。