用分层语义切块解释英语冗余,揭示语言熵随语义复杂度上升的规律。
Semantic Chunking and the Entropy of Natural Language
- 通过自相似语义切块构建多尺度语言结构模型。
- 模型预测的语言熵率与印刷英语实测值一致(约1比特/字符)。
- 适用于研究语言复杂性与信息密度关系的学者。
印刷英语的熵率约为每字符1比特,这一基准值仅在近年被大型语言模型接近。该熵率表明英语相对于随机文本(每字符5比特)具有近80%的冗余。本文提出一种统计模型,旨在捕捉自然语言的多层次结构,从第一性原理解释此冗余水平。模型描述了将文本逐层分割为语义连贯片段直至单个词汇的过程,使文本语义结构得以层级化分解并进行解析处理。基于现代大语言模型与公开数据集的数值实验表明,该模型在不同语义层次上定量刻画了真实文本的结构特征。模型预测的熵率与印刷英语实测熵率相符。此外,理论进一步揭示:自然语言的熵率并非固定,而是随语料语义复杂度系统性升高,该复杂度由模型中唯一的自由参数捕获。
原文摘要 · Abstract (English)
The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. This entropy rate implies that English contains nearly 80 percent redundancy relative to the five bits per character expected for random text. We introduce a statistical model that attempts to capture the intricate multi-scale structure of natural language, providing a first-principles account of this redundancy level. Our model describes a procedure of self-similarly segmenting text into semantically coherent chunks down to the single-word level. The semantic structure of the text can then be hierarchically decomposed, allowing for analytical treatment. Numerical experiments with modern LLMs and open datasets suggest that our model quantitatively captures the structure of real texts at different levels of the semantic hierarchy. The entropy rate predicted by our model agrees with the estimated entropy rate of printed English. Moreover, our theory further reveals that the entropy rate of natural language is not fixed but should increase systematically with the semantic complexity of corpora, which are captured by the only free parameter in our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。