揭示无限宽两层神经网络的精确可存储性相变边界。
Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks
- 采用全三级对称破缺理论计算出精确的可存储与不可存储相变点。
- 发现负感知机在特定参数下典型状态重叠分布存在间隙,导致算法失效。
- 证明梯度下降无法达到理论最大存储容量,适用于理解优化偏差。
我们研究了使用两类连续非凸权重模型存储随机模式-标签关联的问题,包括带负边距的感知机和具有非重叠感受野及通用激活函数的无限宽两层神经网络。通过全三级对称破缺(full-RSB)假设,我们精确计算了SAT/UNSAT相变点。此外,在负感知机情形下,我们发现典型状态的重叠分布在其相图的某些区域呈现重叠间隙(不连通支持),这意味着一些保证近似消息传递(AMP)算法收敛至容量的近期定理不再适用。最后,我们证明梯度下降即使在无重叠间隙时也无法达到最大容量,这与二值权重模型中的现象类似,表明基于梯度的算法倾向于高度非典型的态,而这些态的不可达性决定了算法阈值。
原文摘要 · Abstract (English)
We analyze the problem of storing random pattern-label associations using two classes of continuous non-convex weights models, namely the perceptron with negative margin and an infinite-width two-layer neural network with non-overlapping receptive fields and generic activation function. Using a full-RSB ansatz we compute the exact value of the SAT/UNSAT transition. Furthermore, in the case of the negative perceptron we show that the overlap distribution of typical states displays an overlap gap (a disconnected support) in certain regions of the phase diagram defined by the value of the margin and the density of patterns to be stored. This implies that some recent theorems that ensure convergence of Approximate Message Passing (AMP) based algorithms to capacity are not applicable. Finally, we show that Gradient Descent is not able to reach the maximal capacity, irrespectively of the presence of an overlap gap for typical states. This finding, similarly to what occurs in binary weight models, suggests that gradient-based algorithms are biased towards highly atypical states, whose inaccessibility determines the algorithmic threshold.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。