arXiv:2410.05807cs.LGcs.DS2024-10

拓展凸性和光滑性理论,解释深度学习优化机制。

Extended convexity and smoothness and their applications in deep learning

  • 用扩展的凸性与光滑性分析非凸优化问题。
  • 证明经验风险最小化等价于优化梯度范数与结构误差。
  • 适合研究深度学习优化原理的学者参考。

经典假设如强凸性和Lipschitz光滑性难以刻画深度学习优化问题的本质,后者通常为非凸非光滑,导致传统分析失效。本文通过拓展强凸性和Lipschitz光滑性的概念,揭示深度学习非凸优化的机制。在设定约束下,我们证明经验风险最小化等价于优化局部梯度范数与结构误差,二者共同构成经验风险的上下界。进一步分析表明,随机梯度下降(SGD)可有效降低局部梯度范数;而跳跃连接、过参数化和随机初始化等技术有助于控制结构误差。通过大量实验验证了核心结论。理论与实证结果表明,本研究为理解深度学习中的非凸优化机制提供了新视角。

原文摘要 · Abstract (English)

Classical assumptions like strong convexity and Lipschitz smoothness often fail to capture the nature of deep learning optimization problems, which are typically non-convex and non-smooth, making traditional analyses less applicable. This study aims to elucidate the mechanisms of non-convex optimization in deep learning by extending the conventional notions of strong convexity and Lipschitz smoothness. By leveraging these concepts, we prove that, under the established constraints, the empirical risk minimization problem is equivalent to optimizing the local gradient norm and structural error, which together constitute the upper and lower bounds of the empirical risk. Furthermore, our analysis demonstrates that the stochastic gradient descent (SGD) algorithm can effectively minimize the local gradient norm. Additionally, techniques like skip connections, over-parameterization, and random parameter initialization are shown to help control the structural error. Ultimately, we validate the core conclusions of this paper through extensive experiments. Theoretical analysis and experimental results indicate that our findings provide new insights into the mechanisms of non-convex optimization in deep learning.

优化理论深度学习非凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。