arXiv:2505.18244cs.CLcs.AI2025-05

发现大模型有分层结构,架构决定信息处理方式而非规模大小。

Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation

  • 用信息瓶颈理论建模分层处理机制,揭示模型内部功能边界。
  • 不同架构模型的边界位置差异显著,且与参数规模无关。
  • 跨架构的脆弱性相差千倍,适合研究模型内部结构的学者参考。

为何不同架构家族的语言模型对相同扰动反应迥异?我们认为关键不在于规模,而在于架构如何塑造信息压缩。分析来自 Llama 与 Qwen 家族的八款 Transformer 模型(7B–70B 参数),我们发现每个模型均自发形成离散的功能边界,将层划分为局部、中间和全局处理段——但边界位置及各段脆性主要由架构家族决定,而非模型大小或训练配置。我们提出多尺度概率生成理论(MSPGT),将自回归 Transformer 视为分层变分信息瓶颈系统,并推导出可检验的层级预测。三项预测均被强证实:所有八款模型均出现两个显著相变边界(P1.1);Llama 的边界在十倍参数范围内高度稳定(变异系数 CV=0.067–0.095),而 Qwen 变化剧烈(CV=0.465–0.726),精确匹配强主导与弱主导条件;跨架构局部段脆性相差三个数量级(493倍),该差距仅由架构家族预测,远超家族内或规模驱动的变化。

原文摘要 · Abstract (English)

Why do language models from different architecture families respond so differently to the same perturbation? We argue that the answer is not scale, but \emph{how architecture shapes information compression}. Analyzing eight Transformer models (7B--70B parameters) from the Llama and Qwen families, we show that every model spontaneously develops discrete functional boundaries dividing its layers into Local, Intermediate, and Global processing segments -- yet boundary locations and per-segment brittleness are determined overwhelmingly by architecture family rather than model size or training configuration. We formalize this regularity as the \textbf{Multi-Scale Probabilistic Generation Theory} (MSPGT), which models an autoregressive Transformer as a Hierarchical Variational Information Bottleneck system and derives a tiered set of falsifiable predictions. Three predictions are strongly confirmed: all eight models exhibit two prominent phase-transition boundaries (P1.1); Llama boundary positions are stable across a $10{\times}$ parameter range ($\mathrm{CV}{=}0.067$--$0.095$) while Qwen positions vary widely ($\mathrm{CV}{=}0.465$--$0.726$), precisely matching our strong- and weak-dominance conditions; and cross-architecture local-segment brittleness spans \textbf{three orders of magnitude} ($493{\times}$ ratio) -- a gap that architecture family alone predicts and that dwarfs any within-family or scale-driven variation.

大模型分层结构信息瓶颈架构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。