arXiv:2412.00104cs.LGcond-mat.dis-nn2024-12ICLR被引 21

揭示Transformer在上下文学习中从记忆到泛化的机制转变规律

Differential learning kinetics govern the transition from memorization to generalization during in-context learning

  • 通过理论与实验发现记忆与泛化子电路独立学习,速率差异决定转换
  • 提出记忆缩放定律,定量预测网络开始泛化的任务多样性阈值
  • 解释了ICL出现的长尾分布、双峰行为等复杂现象,适合模型机制研究者

Transformer具备上下文学习(ICL)能力:在不更新权重的情况下利用上下文中的新信息。近期研究表明,当模型在足够多样化的任务上训练时,其从记忆到泛化的转变是急剧的。一种解释认为,网络有限的记忆容量促进了泛化。本文通过一个小型Transformer在合成ICL任务上的分析,结合理论与实验,发现记忆和泛化子电路可视为基本独立。这些子电路的学习速率差异决定了从记忆到泛化的转变,而非容量限制。我们揭示了一条记忆缩放定律,用于确定网络开始泛化的任务多样性阈值。该理论定量解释了多种ICL相关现象,包括ICL出现时间的长尾分布、接近阈值时解的双峰行为、上下文与数据分布统计对ICL的影响,以及ICL的瞬态特性。

原文摘要 · Abstract (English)

Transformers exhibit in-context learning (ICL): the ability to use novel information presented in the context without additional weight updates. Recent work shows that ICL emerges when models are trained on a sufficiently diverse set of tasks and the transition from memorization to generalization is sharp with increasing task diversity. One interpretation is that a network's limited capacity to memorize favors generalization. Here, we examine the mechanistic underpinnings of this transition using a small transformer applied to a synthetic ICL task. Using theory and experiment, we show that the sub-circuits that memorize and generalize can be viewed as largely independent. The relative rates at which these sub-circuits learn explains the transition from memorization to generalization, rather than capacity constraints. We uncover a memorization scaling law, which determines the task diversity threshold at which the network generalizes. The theory quantitatively explains a variety of other ICL-related phenomena, including the long-tailed distribution of when ICL is acquired, the bimodal behavior of solutions close to the task diversity threshold, the influence of contextual and data distributional statistics on ICL, and the transient nature of ICL.

上下文学习Transformer机器学习机制泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。