arXiv:2602.17171cs.LGcs.AI2026-02

对比线性与二次注意力在回归任务中的上下文学习能力

In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks

  • 用线性与二次注意力模型实测回归任务的上下文学习表现
  • 线性注意力在低维回归上表现接近二次注意力,但泛化能力较弱
  • 深度增加会削弱线性注意力的性能,适合小规模任务研究

近期研究表明,Transformer和线性注意力模型能在简单函数类(如线性回归)上实现上下文学习(ICL)。本文针对Garg等人提出的经典线性回归任务,实证研究了这两种注意力机制在ICL行为上的差异。我们评估了各架构的学习质量(MSE)、收敛速度及泛化性能,并分析了模型深度对ICL表现的影响。结果表明,在该设置下,线性注意力与二次注意力既有相似之处,也存在明显局限性。

原文摘要 · Abstract (English)

Recent work has demonstrated that transformers and linear attention models can perform in-context learning (ICL) on simple function classes, such as linear regression. In this paper, we empirically study how these two attention mechanisms differ in their ICL behavior on the canonical linear-regression task of Garg et al. We evaluate learning quality (MSE), convergence, and generalization behavior of each architecture. We also analyze how increasing model depth affects ICL performance. Our results illustrate both the similarities and limitations of linear attention relative to quadratic attention in this setting.

注意力机制上下文学习回归任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。