arXiv:2605.07297stat.MLcs.LG2026-05

提出可自适应谱结构的Transformer泛化界,更准确刻画模型泛化能力。

Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers

  • 基于层间谱范数控制,用Schatten范数衡量权重矩阵复杂度
  • 新界在深度和隐藏维度增长时增长更慢,优于传统范数界
  • 适合研究Transformer泛化机理的理论工作者

理解训练后Transformer为何具有良好泛化性能是现代机器学习理论的核心问题。现有基于范数的泛化界虽消除了对隐藏维度的显式多项式依赖,但通常需预先设定固定范数约束,并常呈现对深度的不利指数依赖。本文推导了多层Transformer的谱自适应后验泛化界。在层间谱范数控制下,界以查询-键、值及前馈权重矩阵的层间Schatten量表示。由于Schatten指数无需预先设定,可训练后分别针对每类矩阵与层选择,边界能自适应地根据学习到的奇异值分布,在谱复杂度与维数、深度相关因子间进行权衡。对BERT适配的主导复杂度因子代理的实证比较表明,本方法诱导的代理随深度和隐藏维度增长更缓慢,优于对应范数代理。整体结果从复杂度角度揭示了训练后Transformer的谱结构如何反映在泛化分析中。

原文摘要 · Abstract (English)

Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds for Transformers remove the explicit polynomial dependence on the hidden dimension, they typically impose fixed norm constraints specified a priori and can exhibit unfavorable exponential dependence on depth. In this paper, we derive spectrum-adaptive post hoc generalization bounds for multi-layer Transformers. Under layerwise spectral norm control, the bounds are expressed in terms of layerwise Schatten quantities of the query-key, value, and feedforward weight matrices. Since the Schatten indices need not be fixed a priori and can instead be selected after training, separately for each matrix type and layer, the bounds adaptively trade off spectral complexity against the dimension- and depth-dependent factors according to the learned singular-value profiles. Empirical comparisons of BERT-adapted proxies for the leading complexity factors suggest that the proxies induced by our bounds grow more slowly with depth and hidden dimension than the corresponding norm-based proxies. Overall, our results provide a complexity-based perspective on how the spectral structure of trained Transformers is reflected in generalization analyses.

Transformer泛化界谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。