arXiv:2411.05353cs.LGcs.AI2024-11被引 2

通过调整非线性与数据对称性,可控制神经网络在模运算中的突现学习现象。

Controlling Grokking with Nonlinearity and Data Symmetry

  • 改变激活函数形状及模型深度宽度,调控模运算中的突现学习行为。
  • 增加非线性后,权重投影模式趋于均匀,可用于分解合数模数P。
  • 基于权重熵的度量可反映模型泛化能力,关联最终层神经元局部熵相关性。

本文证明,在神经网络中对模数P的模算术任务进行学习时,突现学习(grokking)行为可通过调整激活函数的特性以及模型的深度和宽度来控制。将最后一层神经网络权重的偶数主成分投影与奇数主成分投影作图,当层数增加、非线性增强时,这些模式变得显著更均匀。该均匀性可用于对非素数模数P进行因式分解。此外,从层权重熵推导出一种衡量网络泛化能力的指标,而网络非线性程度则与最终层神经元权重局部熵之间的相关性有关。

原文摘要 · Abstract (English)

This paper demonstrates that grokking behavior in modular arithmetic with a modulus P in a neural network can be controlled by modifying the profile of the activation function as well as the depth and width of the model. Plotting the even PCA projections of the weights of the last NN layer against their odd projections further yields patterns which become significantly more uniform when the nonlinearity is increased by incrementing the number of layers. These patterns can be employed to factor P when P is nonprime. Finally, a metric for the generalization ability of the network is inferred from the entropy of the layer weights while the degree of nonlinearity is related to correlations between the local entropy of the weights of the neurons in the final layer.

突现学习非线性控制模运算权重分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。