arXiv:2410.11474cs.LGmath.OC2024-10被引 7

解析Transformer如何通过动态转变实现强上下文学习能力

How Transformers Get Rich: Approximation and Dynamics Analysis

  • 从近似与动态两方面分析诱导头机制的实现方式
  • 训练中出现从4元语法到诱导头的突变式转变
  • 适合关注模型内在工作机制的研究者

Transformer展现出卓越的上下文学习能力,但其理论机制仍不清晰。先前研究(Elhage et al., 2021)识别出一种‘丰富’的上下文机制——诱导头,区别于忽略长程依赖的‘懒惰’n-gram模型。本文从近似与动态两个角度分析Transformer如何实现诱导头。在近似分析中,形式化了标准与广义诱导头机制,并考察各子模块的差异化作用;在动态分析中,针对由4-gram与上下文2-gram组成的合成目标进行训练,精确刻画全过程并发现训练过程中存在从懒惰(4-gram)机制向丰富(诱导头)机制的突变。该结果揭示了模型学习中的关键跃迁过程。

原文摘要 · Abstract (English)

Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remains limited. A recent work (Elhage et al., 2021) identified a ``rich'' in-context mechanism known as induction head, contrasting with ``lazy'' $n$-gram models that overlook long-range dependencies. In this work, we provide both approximation and dynamics analyses of how transformers implement induction heads. In the {\em approximation} analysis, we formalize both standard and generalized induction head mechanisms, and examine how transformers can efficiently implement them, with an emphasis on the distinct role of each transformer submodule. For the {\em dynamics} analysis, we study the training dynamics on a synthetic mixed target, composed of a 4-gram and an in-context 2-gram component. This controlled setting allows us to precisely characterize the entire training process and uncover an {\em abrupt transition} from lazy (4-gram) to rich (induction head) mechanisms as training progresses.

Transformer机制分析上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。