arXiv:2512.24969cond-mat.stat-mechcs.CL2025-12被引 2

LLM揭示英语文本存在跨千字符级的长程依赖关系

Large language models and the entropy of English

  • 用大模型分析英文文本,发现上下文越长熵越低,显示远距离依赖存在
  • 在约10^4字符长度下,编码长度仍持续下降,表明字符间有显著长距相关性
  • 适合研究语言统计特性、模型训练机制或物理类比建模的研究者

我们利用大语言模型(LLMs)揭示了来自多种来源的英文文本中的长程结构。在多数情况下,条件熵或编码长度在上下文长度达到约10^4字符时仍持续下降,表明存在直接的跨距离依赖或相互作用。进一步地,我们独立于模型的数据分析也证实了这些分离距离上的小但显著的相关性。编码长度分布显示出随着上下文长度增大,越来越多字符呈现出确定性特征。在模型训练过程中,长距离和短距离的动态表现不同,说明长程结构是逐步学习到的。这些结果为构建语言或大模型的统计物理模型提供了重要约束。

原文摘要 · Abstract (English)

We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. The conditional entropy or code length in many cases continues to decrease with context length at least to $N\sim 10^4$ characters, implying that there are direct dependencies or interactions across these distances. A corollary is that there are small but significant correlations between characters at these separations, as we show from the data independent of models. The distribution of code lengths reveals an emergent certainty about an increasing fraction of characters at large $N$. Over the course of model training, we observe different dynamics at long and short context lengths, suggesting that long-ranged structure is learned only gradually. Our results constrain efforts to build statistical physics models of LLMs or language itself.

语言模型长程依赖信息熵统计物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。