大多数浅层神经网络的最优解附近都呈强凸,收敛更快。
In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods
- 划分参数空间为高效区与冗余区,聚焦高效区分析
- 在高效区几乎所有回归问题中,局部极小值均有强凸邻域
- 适合关注优化性质与收敛速度的研究者
浅层单隐层神经网络在回归任务中,其均方误差损失函数的局部极小值是否具有强凸邻域,直接影响优化器的渐近收敛速度。本文严格分析了使用解析激活函数的此类网络在回归问题中的该性质的普遍性。将参数空间划分为‘高效区’(无法用更少神经元实现的函数)与‘冗余区’(其余参数)。几乎在所有高效区的回归问题中,优化景观仅包含具有强凸邻域的局部极小值。形式上,我们证明:对某些随机选取的回归问题,优化景观在高效区几乎必然为Morse函数。冗余区维度远小于高效区,且在此区域内局部极小值从不孤立。
原文摘要 · Abstract (English)
Whether or not a local minimum of a cost function has a strongly convex neighborhood greatly influences the asymptotic convergence rate of optimizers. In this article, we rigorously analyze the prevalence of this property for the mean squared error induced by shallow, 1-hidden layer neural networks with analytic activation functions when applied to regression problems. The parameter space is divided into two domains: the 'efficient domain' (all parameters for which the respective realization function cannot be generated by a network having a smaller number of neurons) and the 'redundant domain' (the remaining parameters). In almost all regression problems on the efficient domain the optimization landscape only features local minima that are strongly convex. Formally, we will show that for certain randomly picked regression problems the optimization landscape is almost surely a Morse function on the efficient domain. The redundant domain has significantly smaller dimension than the efficient domain and on this domain, potential local minima are never isolated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。