用对偶凸化方法揭示正则化神经网络的损失景观结构
Exploring the loss landscape of regularized neural networks via convex duality
- 将非凸损失问题转为对偶凸问题,分析解集与驻点结构
- 发现最优解拓扑随网络宽度变化存在相变,且可有连续无穷多最优解
- 结果适用于多层、向量输出等复杂架构,适合研究优化性质的学者
我们通过将正则化神经网络的损失问题转化为等价的凸问题并考察其对偶,探讨了该问题的多个方面:驻点结构、最优解连通性、通向任意全局最优的非增损失路径,以及最优解的非唯一性。从具有标量输出的两层网络出发,首先利用对偶刻画凸问题的解集,并进一步刻画所有驻点。基于此,我们发现全局最优解的拓扑结构随网络宽度变化呈现相变现象,并构造出存在连续无穷多最优解的反例。最后,我们证明解集刻画和连通性结果可推广至不同架构,包括两层向量输出网络和平行三层数网络。
原文摘要 · Abstract (English)
We discuss several aspects of the loss landscape of regularized neural networks: the structure of stationary points, connectivity of optimal solutions, path with nonincreasing loss to arbitrary global optimum, and the nonuniqueness of optimal solutions, by casting the problem into an equivalent convex problem and considering its dual. Starting from two-layer neural networks with scalar output, we first characterize the solution set of the convex problem using its dual and further characterize all stationary points. With the characterization, we show that the topology of the global optima goes through a phase transition as the width of the network changes, and construct counterexamples where the problem may have a continuum of optimal solutions. Finally, we show that the solution set characterization and connectivity results can be extended to different architectures, including two-layer vector-valued neural networks and parallel three-layer neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。