理论证明:单层Transformer可线性收敛于高斯混合分类最优解
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
- 在特定假设下,单层Transformer梯度下降可线性收敛
- 训练与测试提示长度足够大时,预测趋近真实标签分布
- 为ICL机制提供首个严格理论分析,适合理论研究者
尽管变换器在实践中展现出强大的上下文学习能力,但对其支持上下文学习的内在机制的理论理解仍处于初级阶段。本文旨在理论上研究变换器在上下文分类任务中的训练动态。我们证明,在特定假设下,针对高斯混合的上下文分类任务,通过梯度下降训练的单层变换器以线性速率收敛到全局最优模型。我们进一步量化了训练和测试提示长度对变换器上下文学习推理误差的影响。结果表明,当训练和测试提示长度足够大时,变换器的预测趋近于标签的真实分布。实验结果验证了理论发现。
原文摘要 · Abstract (English)
Although transformers have demonstrated impressive capabilities for in-context learning (ICL) in practice, theoretical understanding of the underlying mechanism that allows transformers to perform ICL is still in its infancy. This work aims to theoretically study the training dynamics of transformers for in-context classification tasks. We demonstrate that, for in-context classification of Gaussian mixtures under certain assumptions, a single-layer transformer trained via gradient descent converges to a globally optimal model at a linear rate. We further quantify the impact of the training and testing prompt lengths on the ICL inference error of the trained transformer. We show that when the lengths of training and testing prompts are sufficiently large, the prediction of the trained transformer approaches the ground truth distribution of the labels. Experimental results corroborate the theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。