用变分法重新理解浅层网络,让训练变成解线性方程。
Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks
- 将离散训练转为连续变分问题,参数密度在加权Sobolev空间中优化。
- 最优参数密度可直接解线性系统获得,误差率随宽度提升达O(1/N)。
- 突破传统方法局限,为过参数化提供变分框架,适合理论研究者。
尽管神经网络表现优异,其优化原理仍缺乏理论解释,常面临非凸景观与随机启发式问题。本文提出范式转变:将浅层神经网络的离散训练问题替换为具有良好定义的连续变分近似。我们识别出一类在加权Sobolev空间中参数密度上的λ-凸泛函,并证明这些变分问题全局适定、稳定,且具有意外的几乎C³正则性。与现有基于Wasserstein或均场的方法不同,本方法直接获得椭圆正则性与凸分析优势,可证明最优参数密度通过求解单一线性系统得到,无需迭代优化。我们建立了一般化泛化误差控制,相对于正则化参数α,误差率为1/α;有限宽度网络(大小N)以O(1/N)速率逼近连续最优解。该视角弥合了神经正切核(NTK)与特征学习范式间的鸿沟,为过参数化提供了基于变分微积分的严谨框架。
原文摘要 · Abstract (English)
Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem of shallow neural networks with a well-posed continuum variational surrogate. We identify a family of $λ$-convex functionals over parameter densities in weighted Sobolev spaces and prove that these variational problems are globally well-posed, stable, and exhibit unexpected almost $C^3$ regularity. Unlike existing Wasserstein-based or Mean-Field approaches, which often face limited regularity and discretization challenges, our formulation provides direct access to elliptic regularity and convex analysis. This allows us to prove that the optimal parameter density can be obtained by solving a single linear system, bypassing iterative optimization entirely. We establish explicit generalization error controls at a rate of $1/α$ relative to the regularization parameter, and prove that finite-width networks of size $N$ achieve the continuum optimum at an $O(1/N)$ rate. This perspective bridges the gap between the Neural Tangent Kernel (NTK) and feature-learning regimes, providing a principled framework for understanding over-parameterization through the lens of variational calculus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。