arXiv:2604.07715cs.LGmath.OC2026-04

分析单层神经网络学习过程,提出新激活函数并证明收敛性。

Mathematical analysis of one-layer neural network with fixed biases, a new activation function and other observations

  • 研究带固定偏置的单层ReLU网络,用梯度下降优化。
  • 严格证明了$ L^2 $损失下学习过程收敛,且存在谱偏差。
  • 提出FReX激活函数,适用于需要更好泛化的场景。

我们分析了一个具有ReLU激活函数和固定偏置的简单单隐层神经网络,输入输出均为一维。研究了该模型的连续与离散版本,严格证明了在$ L^2 $平方损失函数下,梯度下降过程的收敛性,并揭示了该学习过程中的谱偏差特性。进一步讨论了激活函数应具备的结构与性质,以及特定算子谱与学习过程的关系。基于此,提出一种替代激活函数——全波整流指数函数(FReX),并探讨了使用该函数时梯度下降的收敛性。

原文摘要 · Abstract (English)

We analyze a simple one-hidden-layer neural network with ReLU activation functions and fixed biases, with one-dimensional input and output. We study both continuous and discrete versions of the model, and we rigorously prove the convergence of the learning process with the $L^2$ squared loss function and the gradient descent procedure. We also prove the spectral bias property for this learning process. Several conclusions of this analysis are discussed; in particular, regarding the structure and properties that activation functions should possess, as well as the relationships between the spectrum of certain operators and the learning process. Based on this, we also propose an alternative activation function, the full-wave rectified exponential function (FReX), and we discuss the convergence of the gradient descent with this alternative activation function.

神经网络梯度下降激活函数收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。