arXiv:2502.16075cs.LGmath.OC2025-02ICML被引 10

揭示非齐次深度网络梯度下降的渐近隐式偏差,证明其方向收敛且最大化分类边界。

Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks

  • 从极小经验风险出发,分析梯度下降在非齐次网络中的方向演化规律。
  • 迭代方向收敛,归一化边距几乎单调上升,权重范数趋于无穷。
  • 结果适用于带残差连接和非齐次激活函数的网络,解决开放问题。

我们建立了在指数损失下通用非齐次深度网络梯度下降(GD)的渐近隐式偏差。具体而言,当初始经验风险足够小时(阈值由网络非齐次性度量决定),我们刻画了三个关键性质:首先,由GD迭代产生的归一化边距几乎单调增加;其次,尽管迭代点的范数发散至无穷大,但其方向收敛;最后,该方向极限满足一个边距最大化问题的Karush-Kuhn-Tucker(KKT)条件。此前关于隐式偏差的研究仅限于齐次网络,而我们的结果适用于满足弱近齐次条件的广泛非齐次网络类,包括含残差连接和非齐次激活函数的网络,从而解决了Ji与Telgarsky(2020)提出的开放问题。

原文摘要 · Abstract (English)

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small empirical risk, where the threshold is determined by a measure of the network's non-homogeneity. First, we show that a normalized margin induced by the GD iterates increases nearly monotonically. Second, we prove that while the norm of the GD iterates diverges to infinity, the iterates themselves converge in direction. Finally, we establish that this directional limit satisfies the Karush-Kuhn-Tucker (KKT) conditions of a margin maximization problem. Prior works on implicit bias have focused exclusively on homogeneous networks; in contrast, our results apply to a broad class of non-homogeneous networks satisfying a mild near-homogeneity condition. In particular, our results apply to networks with residual connections and non-homogeneous activation functions, thereby resolving an open problem posed by Ji and Telgarsky (2020).

梯度下降隐式偏差深度学习理论非齐次网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。