提出任务特定提示可分离数据分布的均值与方差,提升上下文学习效果。
Provable Benefits of Task-Specific Prompts for In-context Learning
- 用任务特定提示与单层注意力模型分离条件分布的均值与方差。
- 实验显示提示调优解释均值,上下文学习解释方差,降低优化难度。
- 适合研究上下文学习机制、提示工程与模型泛化能力的读者。
现代语言模型的上下文学习能力推动了对序列模型的数学理解。近期工作表明,线性注意力模型能模拟投影梯度下降,从上下文窗口数据中隐式学习任务向量。本文考虑一种新设置:全局任务分布可划分为若干条件任务分布的并集。我们研究使用任务特定提示和预测头,通过单层注意力模型学习与条件任务分布相关的先验信息。损失曲面分析显示,任务特定提示实现协方差-均值解耦:提示调优解释分布的条件均值,而方差由上下文学习捕捉。引入任务特定头部进一步完全解耦均值与方差估计。该协方差-均值视角同样解释了为何联合训练提示与注意力权重在预训练后能优于微调。
原文摘要 · Abstract (English)
The in-context learning capabilities of modern language models have motivated a deeper mathematical understanding of sequence models. A line of recent work has shown that linear attention models can emulate projected gradient descent iterations to implicitly learn the task vector from the data provided in the context window. In this work, we consider a novel setting where the global task distribution can be partitioned into a union of conditional task distributions. We then examine the use of task-specific prompts and prediction heads for learning the prior information associated with the conditional task distribution using a one-layer attention model. Our results on loss landscape show that task-specific prompts facilitate a covariance-mean decoupling where prompt-tuning explains the conditional mean of the distribution whereas the variance is learned/explained through in-context learning. Incorporating task-specific head further aids this process by entirely decoupling estimation of mean and variance components. This covariance-mean perspective similarly explains how jointly training prompt and attention weights can provably help over fine-tuning after pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。