arXiv:2507.14177cs.LGcs.AI2025-07

解析平滑激活函数两层神经网络的训练解,揭示其内在机制。

Understanding Two-Layer Neural Networks with Smooth Activation Functions

  • 基于泰勒展开与光滑样条构造理论框架
  • 证明任意维度下具有通用逼近能力
  • 适合研究神经网络原理与逼近理论的学者

本文旨在理解使用平滑激活函数(如传统Sigmoid型)的两层神经网络在反向传播算法下的训练解。提出了四个核心机制:泰勒级数展开构建、节点严格偏序、光滑样条实现及光滑连续性约束。证明了该类网络对任意输入维度均具备通用逼近能力,并解释了训练解的形成原理。所提出的全新证明也丰富了逼近论内容。该工作部分揭开了模型解空间的‘黑箱’迷雾。

原文摘要 · Abstract (English)

This paper aims to understand the training solution, which is obtained by the back-propagation algorithm, of two-layer neural networks whose hidden layer is composed of the units with smooth activation functions, including the usual sigmoid type most commonly used before the advent of ReLUs. The mechanism contains four main principles: construction of Taylor series expansions, strict partial order of knots, smooth-spline implementation and smooth-continuity restriction. The universal approximation for arbitrary input dimensionality is proved and the explanation of training solutions is given. Through the principles proposed, the mystery of ``black box'' of the solution space is largely revealed. The new proofs employed also enrich approximation theory.

神经网络逼近理论平滑激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。