提出可自适应谱结构的Transformer泛化界,更准确刻画模型泛化能力。
Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers
- 基于层间谱范数控制,用Schatten范数衡量权重矩阵复杂度
- 新界在深度和隐藏维度增长时增长更慢,优于传统范数界
- 适合研究Transformer泛化机理的理论工作者
理解训练后Transformer为何具有良好泛化性能是现代机器学习理论的核心问题。现有基于范数的泛化界虽消除了对隐藏维度的显式多项式依赖,但通常需预先设定固定范数约束,并常呈现对深度的不利指数依赖。本文推导了多层Transformer的谱自适应后验泛化界。在层间谱范数控制下,界以查询-键、值及前馈权重矩阵的层间Schatten量表示。由于Schatten指数无需预先设定,可训练后分别针对每类矩阵与层选择,边界能自适应地根据学习到的奇异值分布,在谱复杂度与维数、深度相关因子间进行权衡。对BERT适配的主导复杂度因子代理的实证比较表明,本方法诱导的代理随深度和隐藏维度增长更缓慢,优于对应范数代理。整体结果从复杂度角度揭示了训练后Transformer的谱结构如何反映在泛化分析中。
原文摘要 · Abstract (English)
Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds for Transformers remove the explicit polynomial dependence on the hidden dimension, they typically impose fixed norm constraints specified a priori and can exhibit unfavorable exponential dependence on depth. In this paper, we derive spectrum-adaptive post hoc generalization bounds for multi-layer Transformers. Under layerwise spectral norm control, the bounds are expressed in terms of layerwise Schatten quantities of the query-key, value, and feedforward weight matrices. Since the Schatten indices need not be fixed a priori and can instead be selected after training, separately for each matrix type and layer, the bounds adaptively trade off spectral complexity against the dimension- and depth-dependent factors according to the learned singular-value profiles. Empirical comparisons of BERT-adapted proxies for the leading complexity factors suggest that the proxies induced by our bounds grow more slowly with depth and hidden dimension than the corresponding norm-based proxies. Overall, our results provide a complexity-based perspective on how the spectral structure of trained Transformers is reflected in generalization analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。