arXiv:2508.17256cs.LGcs.AI2025-08

用注意力矩阵有效秩解释大模型为何能泛化

Provable Generalization in Overparameterized Neural Nets

  • 以注意力矩阵的有效秩替代参数量衡量模型容量
  • 理论证明泛化误差随样本量下降,符合大模型实测缩放规律
  • 适合关注大模型泛化机制的读者,尤其对注意力机制感兴趣者

深度神经网络通常参数量远超训练样本数,却仍能良好泛化。经典复杂度度量如VC维或PAC-Bayes界在此过参数化情形下常失效,无法解释Transformer等模型的实践成功。本文提出一种针对基于注意力模型的新容量概念,基于注意力矩阵的有效秩。直觉在于:尽管参数量巨大,注意力的函数维度往往低得多。结果表明,该量可导出泛化界,其对样本量的依赖关系与大规模语言模型中观测到的经验缩放律一致,仅差对数因子。虽非完整过参数化学习理论,但表明注意力的谱特性而非原始参数量,可能是理解模型泛化的正确视角。

原文摘要 · Abstract (English)

Deep neural networks often contain far more parameters than training examples, yet they still manage to generalize well in practice. Classical complexity measures such as VC-dimension or PAC-Bayes bounds usually become vacuous in this overparameterized regime, offering little explanation for the empirical success of models like Transformers. In this work, I explore an alternative notion of capacity for attention-based models, based on the effective rank of their attention matrices. The intuition is that, although the parameter count is enormous, the functional dimensionality of attention is often much lower. I show that this quantity leads to a generalization bound whose dependence on sample size matches empirical scaling laws observed in large language models, up to logarithmic factors. While the analysis is not a complete theory of overparameterized learning, it provides evidence that spectral properties of attention, rather than raw parameter counts, may be the right lens for understanding why these models generalize.

神经网络泛化性注意力机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。