arXiv:2505.14808stat.MLcs.LG2025-05被引 6

从低维子空间视角揭示了上下文学习的分布外泛化边界

Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective

  • 用低秩协方差矩阵建模任务,将分布偏移视为子空间夹角
  • 证明当预训练任务在多个子空间时,模型可泛化至零概率区域
  • 适用于理解大模型在未知分布上的推理能力,适合研究ICL机制者

Transformer 模型在上下文学习(ICL)中的出色表现引发了大量研究,但其在分布外(OOD)能否泛化仍缺乏理论解释。本文提出一个最小数学模型,严格推导出 ICL 可泛化的条件。通过参数化为低秩协方差矩阵的线性回归任务,将分布偏移建模为子空间间的夹角,并推导出单层线性注意力模型在所有角度间插值的条件。我们发现,若预训练任务向量来自多个子空间的并集,则模型可泛化至所有夹角——即使测试区域在训练分布中概率为零。反之,若预训练任务服从单一高斯分布,测试风险仍显著依赖于夹角,表明 ICL 无法实现分布外泛化。我们在 GPT-2 上验证了该现象,并扩展至非线性函数类。

原文摘要 · Abstract (English)

The transformer's remarkable ability to perform in-context learning (ICL) has sparked a wide range of studies designed to understand its strengths and limitations. However, a theoretical understanding of when ICL can and cannot generalize beyond its pre-training data still remains unclear. This paper puts forth a minimal mathematical model that provably identifies when ICL can generalize out-of-distribution (OOD). By studying linear regression tasks parameterized with low-rank covariance matrices, we model distribution shifts as varying angles between subspaces and derive conditions under which a single-layer linear attention model interpolates across all angles. We show that if pre-training task vectors are drawn from a union of subspaces, transformers can generalize to all angle shifts--enabling ICL even in regions with zero probability mass in the training distribution. On the other hand, if the pre-training tasks are drawn from a single Gaussian, the test risk shows a non-negligible dependence on the angle, implying that ICL cannot generalize OOD. We empirically show that our results also hold for models such as GPT-2, and present experiments on how our results extend to nonlinear function classes.

上下文学习分布外泛化子空间分析Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。