arXiv:2512.11255cs.LGcs.AI2025-12被引 3

揭示大模型在上下文学习中隐式更新权重的通用机制

A Simple Generalisation of the Implicit Dynamics of In-Context Learning

  • 提出适用于任意位置、任意层的隐式更新通用公式
  • 实证验证了不同位置和层间隐式更新的一致性
  • 为大模型上下文学习理论贴近实际应用提供支持

上下文学习(ICL)指模型在不更新参数的情况下,仅通过输入中的示例即可学习新任务。与以往依赖简化模型和数据设置的理论不同,近期研究发现变换器模块可被抽象为根据上下文隐式更新前馈网络权重(Dherin et al., 2025)。本文对该结果进行了简单推广:适用于(i)除最后一个外的所有序列位置,(ii)除第一层外的任意变换器块,以及(iii)包含层归一化的更真实残差块。我们在简单的上下文线性回归任务上进行了实证验证,并分析了不同标记之间及跨块的隐式更新关系。这些结果使 Dherin 等人(2025)的理论更接近实际应用,具备在大规模模型上验证的潜力。

原文摘要 · Abstract (English)

In-context learning (ICL) refers to the ability of a model to learn new tasks from examples in its input without any parameter updates. In contrast to previous theories of ICL relying on toy models and data settings, recently it has been shown that an abstraction of a transformer block can be seen as implicitly updating the weights of its feedforward network according to the context (Dherin et al., 2025). Here, we provide a simple generalisation of this result for (i) all sequence positions beyond the last, (ii) any transformer block beyond the first, and (iii) more realistic residual blocks including layer normalisation. We empirically verify our theory on simple in-context linear regression tasks and investigate the relationship between the implicit updates related to different tokens within and between blocks. These results help to bring the theory of Dherin et al. (2025) even closer to practice, with potential for validation on large-scale models.

上下文学习变压器隐式更新理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。