arXiv:2607.23397cs.LGmath.ST2026-07

宽神经网络中,可训练输入权重比固定权重更优且无奇异点。

A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks

  • 对比固定与可训练输入权重的三层网络
  • 可训练时泛化误差更小,且参数空间无奇点
  • 揭示宽网络中奇点对学习的关键影响

分层神经网络在人工智能中广泛应用,但其数学性质尚未完全理解。在无限宽极限下,存在两种理论框架:一种假设参数保持接近初始化,将深度学习简化为固定核的核回归;另一种允许参数远离初始化,需优化核本身。本文研究具有有限但大量隐藏单元的三层神经网络,发现训练输入到隐藏层的权重可获得更小的泛化误差,而固定权重时则在参数空间出现奇点。这些结果表明,奇点在宽神经网络中仍起关键作用。

原文摘要 · Abstract (English)

Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning to kernel regression with a fixed kernel by assuming that the parameters remain close to their initialization, whereas the other allows the parameters to move away from their initialization, requiring the kernel itself to be optimized. In this paper, we study a three-layer neural network with a finite but large number of hidden units. We show that training the input-to-hidden weights yields a smaller generalization error than keeping them fixed. Furthermore, the latter setting exhibits singularities in the parameter space, whereas the former does not. These findings indicate that singularities play an essential role even in wide neural networks.

神经网络泛化误差奇点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。