研究注意力机制如何在噪声标签下仍保持良好泛化能力
Benign Overfitting in Token Selection of Attention Mechanism
- 通过信噪比分析注意力的词元选择过程
- 发现即使拟合噪声标签,仍能保持高泛化性能
- 适用于关注模型鲁棒性与理论解释的研究者
注意力机制是Transformer模型的核心组件,对其词元选择机制的理论理解仍处于探索阶段。本文研究在标签存在噪声的分类问题中,注意力机制的训练动态与泛化能力。通过引入信噪比(SNR)的刻画,我们发现注意力机制的词元选择可实现良性过拟合,即在拟合标签噪声的同时仍保持高泛化性能。此外,实验显示模型在初期过拟合后出现延迟的泛化提升。我们在合成数据和真实数据集上进行了实验证明理论分析的有效性。
原文摘要 · Abstract (English)
Attention mechanism is a fundamental component of the transformer model and plays a significant role in its success. However, the theoretical understanding of how attention learns to select tokens is still an emerging area of research. In this work, we study the training dynamics and generalization ability of the attention mechanism under classification problems with label noise. We show that, with the characterization of signal-to-noise ratio (SNR), the token selection of attention mechanism achieves benign overfitting, i.e., maintaining high generalization performance despite fitting label noise. Our work also demonstrates an interesting delayed acquisition of generalization after an initial phase of overfitting. Finally, we provide experiments to support our theoretical analysis using both synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。