发现两层神经网络梯度具低秩结构,揭示其由数据主成分与异常信号主导。
Low Rank Gradients and Where to Find Them
- 在非各向同性数据下,输入权重梯度近似低秩,由两个秩一成分主导。
- 梯度中主成分贡献占比达95%以上,且受激活函数和训练尺度影响显著。
- 适用于理解复杂数据场景下的梯度演化,适合研究模型优化机制的学者。
本文研究两层神经网络训练损失梯度中的低秩结构,放宽了对训练数据与参数的各向同性假设。考虑一种包含主成分和单个奇异分量(spike)的尖峰数据模型,不依赖独立的数据与权重矩阵,并分析了均值场与神经正切核两种标度情形。结果表明,输入权重梯度近似低秩,主要由两个秩一成分构成:一个与主成分数据残差对齐,另一个与输入数据中的秩一异常信号对齐。我们刻画了训练数据特性、标度方式及激活函数如何调控这两个成分的相对权重。此外,还证明标准正则化方法(如权重衰减、输入噪声、雅可比惩罚)能选择性调节这些成分。合成与真实数据实验验证了理论预测的准确性。
原文摘要 · Abstract (English)
This paper investigates low-rank structure in the gradients of the training loss for two-layer neural networks while relaxing the usual isotropy assumptions on the training data and parameters. We consider a spiked data model in which the bulk can be anisotropic and ill-conditioned, we do not require independent data and weight matrices and we also analyze both the mean-field and neural-tangent-kernel scalings. We show that the gradient with respect to the input weights is approximately low rank and is dominated by two rank-one terms: one aligned with the bulk data-residue , and another aligned with the rank one spike in the input data. We characterize how properties of the training data, the scaling regime and the activation function govern the balance between these two components. Additionally, we also demonstrate that standard regularizers, such as weight decay, input noise and Jacobian penalties, also selectively modulate these components. Experiments on synthetic and real data corroborate our theoretical predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。