用大模型量化话语中每句意义的信息量,突破传统字符熵的局限。
Information Theory of Meaningful Communication
- 以句子为单位衡量语言信息,聚焦有意义的语义单元
- 发现叙事中每句传递约1.5比特意义信息,远低于字符级熵值
- 适合关注语言本质、语义压缩与大模型评估的研究者
香农在其开创性论文中将印刷英语视为平稳随机过程,估算其字符熵约为1比特/字符。然而,作为交流工具的语言与印刷形式有本质区别:(i) 信息的基本单位不是字符或单词,而是分句——语言中最短的有意义片段;(ii) 传播的核心是语义内容,而非具体措辞。本文利用近期发展的大型语言模型,首次实现对有意义叙述中信息量的量化,以每分句比特数衡量所传达的意义信息量。
原文摘要 · Abstract (English)
In Shannon's seminal paper, entropy of printed English, treated as a stationary stochastic process, was estimated to be roughly 1 bit per character. However, considered as a means of communication, language differs considerably from its printed form: (i) the units of information are not characters or even words but clauses, i.e. shortest meaningful parts of speech; and (ii) what is transmitted is principally the meaning of what is being said or written, while the precise phrasing that was used to communicate the meaning is typically ignored. In this study, we show that one can leverage recently developed large language models to quantify information communicated in meaningful narratives in terms of bits of meaning per clause.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。