不同激活函数会改变浅层网络的隐式偏差,影响泛化能力。
The Spectral Bias of Shallow Neural Network Learning is Shaped by the Choice of Non-linearity
- 通过重参数化揭示激活函数对隐式偏差的影响机制。
- 理论证明激活函数决定高频率成分的惩罚方式,影响模型选择解的特性。
- 适用于研究模型泛化性、优化路径与激活函数设计的研究者。
尽管经典统计理论预测大规模过参数化神经网络会出现严重过拟合,但现代神经网络仍能良好泛化。这一现象归因于网络的隐式偏差——在众多能正确标记训练数据的解中,倾向于收敛到具有良好泛化能力的解。本文从新视角探究该偏差,聚焦非线性激活函数如何塑造它。首先,提出一种去除连续权重缩放对称性的重参数化方法;其次,在核态(kernel regime)下,利用该重参数化推广近期将浅层网络与Radon变换关联的成果,推导出一大类激活函数对应的隐式偏差显式公式。具体而言,通过Radon变换与傅里叶变换的联系,将核态的归纳偏差解释为最小化一个依赖于激活函数的谱半范数,该半范数惩罚高频分量。最后,在自适应态(adaptive regime)下,我们证明了局部动力学吸引子的存在,促使多个神经元的输入激活函数为零的超平面形成聚类,从而实现神经元响应函数间的对齐。模拟实验验证了上述理论结果。整体工作深化了对过参数化网络泛化能力及其与隐式偏差关系的理解,为设计更高效、鲁棒的模型提供了潜在路径。
原文摘要 · Abstract (English)
Despite classical statistical theory predicting severe overfitting, modern massively overparameterized neural networks still generalize well. This unexpected property is attributed to the network's so-called implicit bias, which describes its propensity to converge to solutions that generalize effectively, among the many possible that correctly label the training data. The aim of our research is to explore this bias from a new perspective, focusing on how non-linear activation functions contribute to shaping it. First, we introduce a reparameterization which removes a continuous weight rescaling symmetry. Second, in the kernel regime, we leverage this reparameterization to generalize recent findings that relate shallow Neural Networks to the Radon transform, deriving an explicit formula for the implicit bias induced by a broad class of activation functions. Specifically, by utilizing the connection between the Radon transform and the Fourier transform, we interpret the kernel regime's inductive bias as minimizing a spectral seminorm that penalizes high-frequency components, in a manner dependent on the activation function. Finally, in the adaptive regime, we demonstrate the existence of local dynamical attractors that facilitate the formation of clusters of hyperplanes where the input to a neuron's activation function is zero, yielding alignment between many neurons' response functions. We confirm these theoretical results with simulations. All together, our work provides a deeper understanding of the mechanisms underlying the generalization capabilities of overparameterized neural networks and its relation with the implicit bias, offering potential pathways for designing more efficient and robust models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。