首次测量乌克兰语熵值,发现每字符信息量上限约1.201比特。
Entropy of Ukrainian

- 通过志愿者预测句子下一字符,估算乌克兰语复杂度。
- 测得乌克兰语熵上界约为1.201比特/字符。
- 结果可对比大模型表现,方法代码公开可用。
在自然语言处理中,语言熵衡量其不可预测性和复杂度。1951年,香农首次通过参与者预测句子下一个字符的方式,估算出英语的熵值。此后,已有针对英语的多项后续研究及对希伯来语的一项研究。然而,至今尚未有人对乌克兰语进行类似实验。本文通过社交媒体招募184名志愿者,采用与英语研究相似的方法,估算乌克兰语的熵值。最终得到熵的上界为 $H_{upper} ough 1.201$ 比特/字符。同时将该结果与当前大型语言模型的表现进行比较,并公开实验方法与代码,讨论主要挑战。
原文摘要 · Abstract (English)
In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the next character in a sentence, he was able to approximate the entropy of the English language. Several follow-up studies by other authors have since been conducted for English, and one for Hebrew. However, to date, Shannon's experiment has never been conducted for Ukrainian. In this paper, we perform this experiment for Ukrainian by recruiting 184 volunteers using social media channels. We rely on techniques used for English to approximate the entropy value of Ukrainian. The final result is an upper bound of $H_{upper}\approx1.201$ bits per character. We compare this to the performance of current Large Language Models. The methods and code used are also documented and published, along with a discussion of the main challenges encountered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。