arXiv:2410.14837cs.LGmath.AG2024-10NeurIPS被引 4

ReLU网络训练受拓扑障碍限制,可能无法达到全局最优。

Topological obstruction to the training of shallow ReLU neural networks

  • ReLU网络梯度流轨迹被约束在特定双曲面乘积上
  • 当输出为标量时,双曲面可有多个连通分支
  • 初始化不同会导致无法抵达全局最优,适合研究优化机制者阅读

研究浅层ReLU神经网络损失曲面几何与优化轨迹的相互作用,是理解复杂网络行为的基础。本文揭示了使用梯度流训练时,浅层ReLU网络损失曲面存在拓扑障碍。由于ReLU函数的齐次性,训练轨迹被限制在依赖参数初始化的双曲面乘积上。当网络输出为单个标量时,这些双曲面可能具有多个连通分支,从而限制了训练过程中可达参数范围。我们解析计算了连通分支的数量,并讨论了通过神经元缩放和置换能否在分支间映射。在此简单设定下,非连通性导致拓扑障碍,根据初始化不同,可能使全局最优不可达。数值实验验证了该结论。

原文摘要 · Abstract (English)

Studying the interplay between the geometry of the loss landscape and the optimization trajectories of simple neural networks is a fundamental step for understanding their behavior in more complex settings. This paper reveals the presence of topological obstruction in the loss landscape of shallow ReLU neural networks trained using gradient flow. We discuss how the homogeneous nature of the ReLU activation function constrains the training trajectories to lie on a product of quadric hypersurfaces whose shape depends on the particular initialization of the network's parameters. When the neural network's output is a single scalar, we prove that these quadrics can have multiple connected components, limiting the set of reachable parameters during training. We analytically compute the number of these components and discuss the possibility of mapping one to the other through neuron rescaling and permutation. In this simple setting, we find that the non-connectedness results in a topological obstruction, which, depending on the initialization, can make the global optimum unreachable. We validate this result with numerical experiments.

神经网络优化障碍拓扑分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。