arXiv:2410.18938stat.MLcs.LG2024-10被引 14

从随机矩阵理论出发,揭示两层神经网络训练后特征谱的演化规律及其对泛化能力的影响。

A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities

  • 通过随机矩阵分析,建立训练后特征与各向同性尖峰模型的等价关系。
  • 推导出特征经验协方差矩阵的确定性等价形式,精确描述谱尾变化。
  • 适用于高学习率、有限初始化的严格理论框架,适合研究深层特征表达能力。

神经网络在训练中适应数据的能力是其核心特性,但当前对其特征学习机制与泛化性能之间关系的数学理解仍不充分。本文基于随机矩阵理论,研究全连接两层神经网络在单次激进梯度下降步后对目标函数的适应过程。在大批次极限下,我们严格证明了更新后的特征等价于各向同性尖峰随机特征模型。针对该模型,我们推导出特征经验协方差矩阵的确定性等价描述,由某些低维算子表示。这使得我们能精确刻画训练对渐近特征谱的影响,特别是谱尾的变化。该确定性等价进一步给出了精确的渐近泛化误差,揭示了特征学习提升泛化性能的内在机制。我们的结果超越了标准随机矩阵系综,具有独立的技术价值。不同于以往工作,本结果适用于极具挑战性的最大学习率情形,完全严格,并允许第二层权重为有限支撑初始化,这对研究学习特征的功能表达力至关重要。该工作提供了超越随机特征与懒惰训练范式的两层网络特征学习泛化机制的精确描述。

原文摘要 · Abstract (English)

A key property of neural networks is their capacity of adapting to data during training. Yet, our current mathematical understanding of feature learning and its relationship to generalization remain limited. In this work, we provide a random matrix analysis of how fully-connected two-layer neural networks adapt to the target function after a single, but aggressive, gradient descent step. We rigorously establish the equivalence between the updated features and an isotropic spiked random feature model, in the limit of large batch size. For the latter model, we derive a deterministic equivalent description of the feature empirical covariance matrix in terms of certain low-dimensional operators. This allows us to sharply characterize the impact of training in the asymptotic feature spectrum, and in particular, provides a theoretical grounding for how the tails of the feature spectrum modify with training. The deterministic equivalent further yields the exact asymptotic generalization error, shedding light on the mechanisms behind its improvement in the presence of feature learning. Our result goes beyond standard random matrix ensembles, and therefore we believe it is of independent technical interest. Different from previous work, our result holds in the challenging maximal learning rate regime, is fully rigorous and allows for finitely supported second layer initialization, which turns out to be crucial for studying the functional expressivity of the learned features. This provides a sharp description of the impact of feature learning in the generalization of two-layer neural networks, beyond the random features and lazy training regimes.

神经网络随机矩阵泛化能力特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。