arXiv:2502.05668cs.LGcs.NE2025-02中稿 · /presented at the …被引 2

揭示了随机子梯度下降在深度网络中的隐式偏置机制

The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks

  • 将归一化SGD看作保守场流的离散化,分析其后期训练动态
  • 证明归一化参数收敛到归一化间隔的临界点集合
  • 首次扩展非光滑随机情形下梯度下降的理论分析

我们分析了常步长随机子梯度下降(SGD)的隐式偏置。研究设定为具有同质神经网络的二分类问题——这是一类包含无偏置MLP和CNN等的深层网络。我们将归一化SGD迭代过程解释为与归一化分类间隔相关的保守场流的欧拉型离散化。基于此解释,我们证明:在训练后期(假设数据已被正归一化间隔正确分类),归一化SGD迭代收敛至归一化间隔的临界点集合。据我们所知,这是首个将Lyu和Li(2020)关于梯度下降离散动力学的分析扩展至非光滑和随机设置的工作。主要结果适用于指数或逻辑损失的二分类任务,并进一步讨论了更一般设置的推广。

原文摘要 · Abstract (English)

We analyze the implicit bias of constant step stochastic subgradient descent (SGD). We consider the setting of binary classification with homogeneous neural networks - a large class of deep neural networks with ReLU-type activation functions such as MLPs and CNNs without biases. We interpret the dynamics of normalized SGD iterates as an Euler-like discretization of a conservative field flow that is naturally associated to the normalized classification margin. Owing to this interpretation, we show that normalized SGD iterates converge to the set of critical points of the normalized margin at late-stage training (i.e., assuming that the data is correctly classified with positive normalized margin). Up to our knowledge, this is the first extension of the analysis of Lyu and Li (2020) on the discrete dynamics of gradient descent to the nonsmooth and stochastic setting. Our main result applies to binary classification with exponential or logistic losses. We additionally discuss extensions to more general settings.

优化分析深度学习随机梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。