arXiv:2501.17745cs.LG2025-01被引 14

揭示Transformer在上下文线性回归中从通用到专用的动态演化过程

Dynamics of Transient Structure in In-Context Linear Regression Transformers

  • 通过轨迹主成分分析发现模型先表现如岭回归
  • 训练中期出现过渡态,最终适配具体任务分布
  • 用局部学习系数验证复杂度与损失权衡机制

现代深度神经网络展现出丰富的内部计算结构。理解这些结构的形成机制是深度学习科学的核心目标。本文研究了在具有中间任务多样性的上下文线性回归任务中,Transformer表现出的瞬态脊现象:模型在训练初期行为类似岭回归,随后逐渐专业化为适应其训练分布中的具体任务。这一从通用解到专用解的转变通过联合轨迹主成分分析得以揭示。进一步,基于贝叶斯内模型选择理论,我们提出一个普遍解释:瞬态结构源于损失与复杂度之间动态权衡。我们通过测量变压器的局部学习系数来定义模型复杂度,实证验证了该解释。

原文摘要 · Abstract (English)

Modern deep neural networks display striking examples of rich internal computational structure. Uncovering principles governing the development of such structure is a priority for the science of deep learning. In this paper, we explore the transient ridge phenomenon: when transformers are trained on in-context linear regression tasks with intermediate task diversity, they initially behave like ridge regression before specializing to the tasks in their training distribution. This transition from a general solution to a specialized solution is revealed by joint trajectory principal component analysis. Further, we draw on the theory of Bayesian internal model selection to suggest a general explanation for the phenomena of transient structure in transformers, based on an evolving tradeoff between loss and complexity. We empirically validate this explanation by measuring the model complexity of our transformers as defined by the local learning coefficient.

Transformer动态结构学习机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。