arXiv:2505.15013stat.MLcs.LG2025-05被引 1

首次证明Adam在深度ReLU网络中的收敛性与泛化能力。

Convergence of Adam in Deep ReLU Networks via Directional Complexity and Kakeya Bounds

  • 基于分层莫尔斯理论与Kakeya集新结果,分析区域穿越次数
  • 区域穿越数从指数级降至近线性,突破传统平滑假设限制
  • 适合研究优化理论的学者,尤其关注自适应方法的泛化性

一阶自适应优化方法如Adam是训练现代深度神经网络的首选。尽管其在实践中表现优异,但在非光滑场景(尤其是深度ReLU网络)中的理论理解仍不充分。ReLU激活函数导致指数级的区域边界,使标准光滑性假设失效。本文首次建立了Adam在深度ReLU网络中的\(\tilde{O}\bigl(\sqrt{d_{\mathrm{eff}}/n}\bigr)\)泛化界,并在无全局PL或凸性假设下实现了非光滑、非凸ReLU景观下的全局最优收敛。分析基于分层莫尔斯理论和新的Kakeya集成果,提出多层精炼框架,逐步收紧区域穿越的上界。证明区域穿越数由指数级降至近线性。借助基于Kakeya的方法,所得泛化界优于PAC-Bayes方法,并在温和的统一低屏障假设下展示收敛性。

原文摘要 · Abstract (English)

First-order adaptive optimization methods like Adam are the default choices for training modern deep neural networks. Despite their empirical success, the theoretical understanding of these methods in non-smooth settings, particularly in Deep ReLU networks, remains limited. ReLU activations create exponentially many region boundaries where standard smoothness assumptions break down. \textbf{We derive the first \(\tilde{O}\!\bigl(\sqrt{d_{\mathrm{eff}}/n}\bigr)\) generalization bound for Adam in Deep ReLU networks and the first global-optimal convergence for Adam in the non smooth, non convex relu landscape without a global PL or convexity assumption.} Our analysis is based on stratified Morse theory and novel results in Kakeya sets. We develop a multi-layer refinement framework that progressively tightens bounds on region crossings. We prove that the number of region crossings collapses from exponential to near-linear in the effective dimension. Using a Kakeya based method, we give a tighter generalization bound than PAC-Bayes approaches and showcase convergence using a mild uniform low barrier assumption.

优化理论深度学习AdamReLU网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。