arXiv:2510.04285cs.CLcond-mat.stat-mech2025-10被引 1

用累积量分析语言模型如何捕捉高阶统计结构。

Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy

  • 将softmax熵视为扰动,推导出可分离高阶相关性的闭式累积量。
  • 训练中累积量单调上升并饱和,反映模型从学方差到学偏度、峰度的过程。
  • 数学文本与普通文本有不同累积量特征,揭示不同处理机制。

我们提出一种累积量展开框架,用于量化大语言模型在下一个词预测过程中对高阶统计结构的内化程度。通过将每一层的logit分布的softmax熵视为相对于其“中心”分布的扰动,推导出可分离逐级高阶相关性的闭式累积量可观测值。在Pile-10K提示上对GPT-2和Pythia模型进行实证分析发现:(i) 结构化提示呈现典型的先上升后平缓的层间变化趋势,而随机打乱的提示则保持平坦,表明累积量轮廓依赖于有意义的上下文;(ii) 训练过程中所有累积量均单调递增并趋于饱和,直接可视化了模型从捕捉方差到学习偏度、峰度及更高阶统计结构的演进过程;(iii) 数学类提示表现出与一般文本显著不同的累积量特征,定量揭示模型对数学内容与语言内容采用根本不同的处理机制。这些结果确立了累积量分析作为高维神经网络特征学习动态的一种轻量、数学严谨的探测工具。

原文摘要 · Abstract (English)

We introduce a cumulant-expansion framework for quantifying how large language models (LLMs) internalize higher-order statistical structure during next-token prediction. By treating the softmax entropy of each layer's logit distribution as a perturbation around its "center" distribution, we derive closed-form cumulant observables that isolate successively higher-order correlations. Empirically, we track these cumulants in GPT-2 and Pythia models on Pile-10K prompts. (i) Structured prompts exhibit a characteristic rise-and-plateau profile across layers, whereas token-shuffled prompts remain flat, revealing the dependence of the cumulant profile on meaningful context. (ii) During training, all cumulants increase monotonically before saturating, directly visualizing the model's progression from capturing variance to learning skew, kurtosis, and higher-order statistical structures. (iii) Mathematical prompts show distinct cumulant signatures compared to general text, quantifying how models employ fundamentally different processing mechanisms for mathematical versus linguistic content. Together, these results establish cumulant analysis as a lightweight, mathematically grounded probe of feature-learning dynamics in high-dimensional neural networks.

语言模型统计结构累积量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。