SGD在宽网络中导致有效宽度坍缩,解具有有限结构。
Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

- 通过均值场分析揭示训练动态的极限行为
- 解函数为分段仿射,超平面数不超过2P-1(P为数据可分类型数)
- 学习到的特征方向具唯一激活模式,适合研究深层网络泛化
我们研究了噪声随机梯度下降在训练宽两层ReLU网络进行多元回归时的隐式偏差。在均值场框架下,训练动态近似为一个Wasserstein梯度流,收敛至唯一驻定测度。我们刻画了该驻定测度及其对应的预测器结构:尽管网络无限过参数化,所学预测器仍具有有效有限表示——输入权重和偏置沿有限多个方向对齐,引发有效宽度坍缩。具体而言,解函数是连续分段仿射的,其仿射区域由有限个超平面排列决定。学习方向数(即超平面数)上界为 $2 ilde{P}-1$,其中 $ ilde{P}$ 表示训练输入可实现的线性二分类型数。我们进一步证明了学习表示的非冗余性:每个学习方向在训练数据上诱导出唯一的三元激活模式。因此,所学预测器的复杂度由训练数据的组合几何决定。
原文摘要 · Abstract (English)
We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression. In a mean-field regime, the training dynamics are approximated by a Wasserstein gradient flow that converges to a unique stationary measure. We characterize the structure of this stationary measure and the predictor it represents. We show that, despite the network being infinitely overparameterized, the learned predictor admits an effectively finite representation: the input weights and biases align along finitely many directions, leading to an effective width collapse. In particular, the solution function is continuous piecewise affine, with affine regions determined by the cells of a finite hyperplane arrangement. The number of learned directions, and hence hyperplanes, is bounded above by $2\mathcal{P}-1$, where $\mathcal{P}$ denotes the number of linear dichotomies realizable on the training inputs. We further establish a non-redundancy property of the learned representation by proving that each learned direction induces a unique ternary activation pattern on the training data. Consequently, the complexity of the learned predictor is governed by the combinatorial geometry of the training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。