arXiv:2603.21541cs.LGcs.AI2026-03被引 1

提出更精确的Transformer泛化误差上界,揭示架构差异对泛化能力的影响。

Sharper Generalization Bounds for Transformer

  • 用偏置Rademacher复杂度建模Transformer泛化误差
  • 基于矩阵秩和范数控制假设空间覆盖数,获得最优收敛率上界
  • 放宽特征有界假设,适用于子高斯和重尾分布场景

本文研究Transformer模型的泛化误差上界。基于偏置Rademacher复杂度,推导出单层单头、单层多头及多层Transformer的不同泛化上界。首先将Transformer的过拟合风险表示为偏置Rademacher复杂度的形式;通过其与假设空间经验覆盖数的关联,得到达到最优收敛速率(常数因子内)的过拟合风险上界。随后,利用矩阵秩和矩阵范数对Transformer假设空间的覆盖数进行上界估计,获得精确且依赖于架构的泛化上界。最后,放宽特征映射有界性假设,将理论结果推广至子高斯特征和重尾分布设置。

原文摘要 · Abstract (English)

This paper studies generalization error bounds for Transformer models. Based on the offset Rademacher complexity, we derive sharper generalization bounds for different Transformer architectures, including single-layer single-head, single-layer multi-head, and multi-layer Transformers. We first express the excess risk of Transformers in terms of the offset Rademacher complexity. By exploiting its connection with the empirical covering numbers of the corresponding hypothesis spaces, we obtain excess risk bounds that achieve optimal convergence rates up to constant factors. We then derive refined excess risk bounds by upper bounding the covering numbers of Transformer hypothesis spaces using matrix ranks and matrix norms, leading to precise, architecture-dependent generalization bounds. Finally, we relax the boundedness assumption on feature mappings and extend our theoretical results to settings with unbounded (sub-Gaussian) features and heavy-tailed distributions.

Transformer泛化界理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。