arXiv:2509.15152stat.MLcs.LG2025-09被引 3

揭示随机Transformer在上下文学习中的等效模型规律

Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models

  • 用随机初始化的Transformer研究上下文学习,第二层可训练
  • 理论证明其误差表现等价于有限阶赫尔米特多项式模型
  • 验证了非线性、过参数化对学习能力的影响,适合理论研究者

我们研究预训练Transformer在非线性回归任务中进行上下文学习(ICL)的能力。重点关注一个带有非线性MLP头的随机Transformer,其中第一层随机初始化并固定,第二层可训练。在上下文长度、输入维度、隐藏维度、训练任务数和训练样本数共同趋于无穷的渐近设定下,我们证明该随机Transformer在ICL误差上等价于一个有限阶赫尔米特多项式模型。这一等价关系通过多种激活函数、上下文长度、隐层宽度(揭示双下降现象)及正则化设置下的仿真得到验证。结果为MLP层如何提升ICL提供了理论与实证洞察,并揭示了非线性与过参数化对模型性能的影响。

原文摘要 · Abstract (English)

We study the in-context learning (ICL) capabilities of pretrained Transformers in the setting of nonlinear regression. Specifically, we focus on a random Transformer with a nonlinear MLP head where the first layer is randomly initialized and fixed while the second layer is trained. Furthermore, we consider an asymptotic regime where the context length, input dimension, hidden dimension, number of training tasks, and number of training samples jointly grow. In this setting, we show that the random Transformer behaves equivalent to a finite-degree Hermite polynomial model in terms of ICL error. This equivalence is validated through simulations across varying activation functions, context lengths, hidden layer widths (revealing a double-descent phenomenon), and regularization settings. Our results offer theoretical and empirical insights into when and how MLP layers enhance ICL, and how nonlinearity and over-parameterization influence model performance.

Transformer上下文学习理论分析非线性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。