arXiv:2502.14114cs.LGcs.AI2025-02被引 2

揭示过参数深度网络达零损失的条件并给出显式解法

Zero loss guarantees and explicit minimizers for generic overparametrized Deep Learning networks

  • 通过分析提出过参数网络实现零损失的充分条件
  • 构造出无需梯度下降的零损失最小化器
  • 指出深层网络可能降低梯度下降优化效率

我们确定了在监督学习框架下,对于$\/mathcal{L}^2$损失和通用训练数据,过参数化深度学习(DL)网络达到零损失的充分条件。本文提出了一个显式构造零损失最小化器的方法,无需依赖梯度下降。另一方面,通过分析训练雅可比矩阵秩损失的条件,指出增加网络深度可能降低梯度下降算法的成本最小化效率。研究结果阐明了欠参数与过参数深度学习中零损失可达性之间的关键差异。

原文摘要 · Abstract (English)

We determine sufficient conditions for overparametrized deep learning (DL) networks to guarantee the attainability of zero loss in the context of supervised learning, for the $\mathcal{L}^2$ cost and {\em generic} training data. We present an explicit construction of the zero loss minimizers without invoking gradient descent. On the other hand, we point out that increase of depth can deteriorate the efficiency of cost minimization using a gradient descent algorithm by analyzing the conditions for rank loss of the training Jacobian. Our results clarify key aspects on the dichotomy between zero loss reachability in underparametrized versus overparametrized DL.

深度学习过参数化零损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。