arXiv:2501.07359cs.CLcs.AI2025-01被引 2

大模型层间语义层级并非线性,深层反而会压缩早期信息并出现异常波动。

Emergent effects of scaling on the functional hierarchies within large language models

  • 用支持向量机和岭回归分析各层激活,发现语义表征随层数变化
  • 小模型中语义逐步深化,大模型中深层表征出现先降后升的奇异波动
  • 深层不仅抽象全局信息,还无意义压缩早期上下文,且相邻层协同分工

本研究通过输入简单文本(如"A church and organ")并提取大型语言模型(如Llama-3.2-3b,28层)各层激活值,使用支持向量机与岭回归预测文本标签,检验每层是否编码特定信息。结果部分支持传统功能层级观点:物品语义在浅层(2-7层)最强,两物关系在中层(8-12层),四物类比在更深层(10-15层)。但深层随后逐渐减少对局部信息的表示,转而关注全局抽象。然而,大模型(Llama-3.3-70b-Instruct)表现出反常现象:两物关系与四物类比的表征随深度增加先上升后骤降,再短暂回升,此模式在多个实验中重复出现。此外,模型规模增大还引发相邻层注意力机制的动态协调,表现为层间分工交替波动。表明尽管存在抽象层级,大规模模型仍呈现复杂非线性特征。

原文摘要 · Abstract (English)

Large language model (LLM) architectures are often described as functionally hierarchical: Early layers process syntax, middle layers begin to parse semantics, and late layers integrate information. The present work revisits these ideas. This research submits simple texts to an LLM (e.g., "A church and organ") and extracts the resulting activations. Then, for each layer, support vector machines and ridge regressions are fit to predict a text's label and thus examine whether a given layer encodes some information. Analyses using a small model (Llama-3.2-3b; 28 layers) partly bolster the common hierarchical perspective: Item-level semantics are most strongly represented early (layers 2-7), then two-item relations (layers 8-12), and then four-item analogies (layers 10-15). Afterward, the representation of items and simple relations gradually decreases in deeper layers that focus on more global information. However, several findings run counter to a steady hierarchy view: First, although deep layers can represent document-wide abstractions, deep layers also compress information from early portions of the context window without meaningful abstraction. Second, when examining a larger model (Llama-3.3-70b-Instruct), stark fluctuations in abstraction level appear: As depth increases, two-item relations and four-item analogies initially increase in their representation, then markedly decrease, and afterward increase again momentarily. This peculiar pattern consistently emerges across several experiments. Third, another emergent effect of scaling is coordination between the attention mechanisms of adjacent layers. Across multiple experiments using the larger model, adjacent layers fluctuate between what information they each specialize in representing. In sum, an abstraction hierarchy often manifests across layers, but large models also deviate from this structure in curious ways.

大模型架构层级表征注意力机制模型缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。