arXiv:2510.04548cond-mat.dis-nncs.LG2025-10中稿 · AISTATS 2026被引 2

解析Transformer如何通过低秩任务结构实现上下文学习

Learning Linear Regression with Low-Rank Tasks in-Context

  • 用低秩回归任务分析线性注意力模型的上下文学习机制
  • 发现有限预训练数据引发隐式正则化,影响泛化误差
  • 揭示任务结构决定泛化误差的尖锐相变现象

上下文学习(ICL)是现代大语言模型的核心组件,但其理论机制仍不清晰。本文针对真实场景中任务具有共同结构的情况,研究一个在低秩回归任务上训练的线性注意力模型。在高维极限下,精确刻画了预测分布与泛化误差。结果表明,有限预训练数据引起的统计波动会诱导隐式正则化,并发现泛化误差存在由任务结构决定的尖锐相变。这些发现为理解Transformer如何学会任务结构提供了理论框架。

原文摘要 · Abstract (English)

In-context learning (ICL) is a key building block of modern large language models, yet its theoretical mechanisms remain poorly understood. It is particularly mysterious how ICL operates in real-world applications where tasks have a common structure. In this work, we address this problem by analyzing a linear attention model trained on low-rank regression tasks. Within this setting, we precisely characterize the distribution of predictions and the generalization error in the high-dimensional limit. Moreover, we find that statistical fluctuations in finite pre-training data induce an implicit regularization. Finally, we identify a sharp phase transition of the generalization error governed by task structure. These results provide a framework for understanding how transformers learn to learn the task structure.

上下文学习线性注意力低秩任务泛化误差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。