arXiv:2502.12508cs.LG2025-02被引 3

解析Transformer泛化能力,揭示训练动态与过拟合的关系。

Understanding Generalization in Transformers: Error Bounds and Training Dynamics Under Benign and Harmful Overfitting

  • 构建两层Transformer的泛化理论,分析信号噪声比下的误差边界。
  • 发现训练分三个阶段,不同阶段对应不同泛化误差范围。
  • 实验验证理论预测,适合研究模型泛化机制的学者参考。

Transformer作为众多大规模模型的基础架构,表现出在训练数据上过拟合但仍能保持对未见数据的良好泛化能力,这一现象称为良性过拟合。然而,关于训练动态如何影响良性过拟合背景下误差边界的研究所限。本文针对带标签翻转噪声的两层Transformer,建立泛化理论,给出了在不同信噪比(SNR)下良性和有害过拟合的泛化误差边界。训练动态被划分为三个不同阶段,每个阶段对应特定的误差边界。此外,通过大量实验识别影响Transformer测试误差的关键因素,实验结果与理论预测高度一致,验证了结论的有效性。

原文摘要 · Abstract (English)

Transformers serve as the foundational architecture for many successful large-scale models, demonstrating the ability to overfit the training data while maintaining strong generalization on unseen data, a phenomenon known as benign overfitting. However, research on how the training dynamics influence error bounds within the context of benign overfitting has been limited. This paper addresses this gap by developing a generalization theory for a two-layer transformer with labeled flip noise. Specifically, we present generalization error bounds for both benign and harmful overfitting under varying signal-to-noise ratios (SNR), where the training dynamics are categorized into three distinct stages, each with its corresponding error bounds. Additionally, we conduct extensive experiments to identify key factors that influence test errors in transformers. Our experimental results align closely with the theoretical predictions, validating our findings.

Transformer泛化理论过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。