arXiv:2603.12541cs.LGcs.SY2026-03

大模型越复杂,其深层变化越趋线性,可用简单模型精准模拟。

As Language Models Scale, Low-order Linear Depth Dynamics Emerge

  • 用32维线性模型精确复现GPT-2-large各层敏感度分布
  • 模型越大,线性逼近效果越好,呈单调提升趋势
  • 可实现低能耗多层干预,优于传统启发式方法

大型语言模型常被视为高维非线性系统并被当作黑箱处理。本文发现,在上下文范围内,Transformer的深度动态可被准确的低阶线性代理模型捕捉。在毒性检测、讽刺识别、仇恨言论和情感分析等任务中,一个32维线性代理模型与GPT-2-large的逐层敏感度分布几乎完全一致,能精确反映在每层添加扰动时最终输出的变化。进一步揭示了一个令人惊讶的缩放规律:对于固定阶数的线性代理模型,其与完整模型的一致性随模型规模增大而单调提升,覆盖整个GPT-2系列。该线性代理还支持可解释的多层干预策略,其能耗低于标准启发式调度方案。综合来看,随着语言模型规模扩大,上下文内的低阶线性深度动态逐渐显现,为分析与控制模型提供了系统理论基础。

原文摘要 · Abstract (English)

Large language models are often viewed as high-dimensional nonlinear systems and treated as black boxes. Here, we show that transformer depth dynamics admit accurate low-order linear surrogates within context. Across tasks including toxicity, irony, hate speech and sentiment, a 32-dimensional linear surrogate reproduces the layerwise sensitivity profile of GPT-2-large with near-perfect agreement, capturing how the final output shifts under additive injections at each layer. We then uncover a surprising scaling principle: for a fixed-order linear surrogate, agreement with the full model improves monotonically with model size across the GPT-2 family. This linear surrogate also enables principled multi-layer interventions that require less energy than standard heuristic schedules when applied to the full model. Together, our results reveal that as language models scale, low-order linear depth dynamics emerge within contexts, offering a systems-theoretic foundation for analyzing and controlling them.

模型分析线性近似深度动态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。