揭示预训练与测试任务对齐程度如何决定上下文学习的泛化能力
Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
- 提出新度量指标,量化预训练任务信息对测试时推理的帮助程度
- 发现任务分布对齐不足会显著降低上下文学习性能,且存在专精与泛化权衡
- 理论结果在非线性Transformer中验证,适用于多种实际场景
上下文学习(ICL)是Transformer模型的核心能力,但其出现与鲁棒性所依赖的数据结构仍不明确。本文研究预训练任务结构如何影响ICL的泛化表现。通过一个可解的线性回归上下文学习模型,我们推导出高维情况下任意预训练-测试任务协方差不匹配下的ICL泛化误差精确表达式。由此提出一种新的对齐度量,衡量预训练任务分布信息在测试时的有用性。该度量不仅能准确预测可解模型中的ICL性能,也在非线性Transformer中得到验证。分析进一步揭示:在任务分布对齐程度不同的情况下,增加预训练任务多样性可能提升或损害测试表现,存在专精与泛化之间的权衡。这些结果表明,训练-测试任务对齐是决定ICL泛化能力的关键因素。
原文摘要 · Abstract (English)
In-context learning (ICL) is a central capability of Transformer models, but the structures in data that enable its emergence and govern its robustness remain poorly understood. In this work, we study how the structure of pretraining tasks governs generalization in ICL. Using a solvable model for ICL of linear regression by linear attention, we derive an exact expression for ICL generalization error in high dimensions under arbitrary pretraining-testing task covariance mismatch. This leads to a new alignment measure that quantifies how much information about the pretraining task distribution is useful for inference at test time. We show that this measure directly predicts ICL performance not only in the solvable model but also in nonlinear Transformers. Our analysis further reveals a tradeoff between specialization and generalization in ICL: depending on task distribution alignment, increasing pretraining task diversity can either improve or harm test performance. Together, these results identify train-test task alignment as a key determinant of generalization in ICL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。