arXiv:2411.07071cs.LGcs.AI2024-11被引 1

通过微扰分析揭示大模型中归纳推理的普遍涌现机制

Universal Response and Emergence of Induction in LLMs

  • 用残差流单令牌微扰探测模型响应,发现尺度不变性
  • 在中间层逐步观测到诱导行为的信号,跨模型一致
  • 为大模型电路分析提供可量化的基准方法

尽管归纳被视为大语言模型上下文学习的关键机制,但其在真实模型中的精确电路分解仍不清晰。本文通过探测残差流的弱单令牌微扰响应,研究大模型中归纳行为的涌现。我们发现模型在响应上存在稳健且普适的尺度不变性,使我们能够量化整个模型中令牌相关性的积累过程。应用该方法后,在Gemma-2-2B、Llama-3.2-3B和GPT-2-XL的残差流中均观测到归纳行为的特征信号。所有模型中,这些信号均在中间层逐渐出现,并识别出构成该行为的关键模型组件。结果揭示了大模型内部组件的协同作用,为大规模电路分析提供了基准。

原文摘要 · Abstract (English)

While induction is considered a key mechanism for in-context learning in LLMs, understanding its precise circuit decomposition beyond toy models remains elusive. Here, we study the emergence of induction behavior within LLMs by probing their response to weak single-token perturbations of the residual stream. We find that LLMs exhibit a robust, universal regime in which their response remains scale-invariant under changes in perturbation strength, thereby allowing us to quantify the build-up of token correlations throughout the model. By applying our method, we observe signatures of induction behavior within the residual stream of Gemma-2-2B, Llama-3.2-3B, and GPT-2-XL. Across all models, we find that these induction signatures gradually emerge within intermediate layers and identify the relevant model sections composing this behavior. Our results provide insights into the collective interplay of components within LLMs and serve as a benchmark for large-scale circuit analysis.

大模型机制归纳推理残差流分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。