深度网络用梯度下降学高维分层函数,样本需求远低于浅层网络。
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
- 设计分层高斯目标函数,分析深度与浅层网络的学习机制差异。
- 深度网络通过梯度下降逐步降低有效维度,仅需少量样本即可学习。
- 揭示深度模型本质优势,适合研究深度学习理论的学者参考。
理解深度神经网络在梯度下降(GD)训练下相较于浅层模型的计算优势,仍是未解的理论难题。本文引入一类目标函数(单/多指数高斯分层目标),包含潜在子空间的层级维度结构。该框架使我们能在高维极限下,解析分析深度网络与浅层网络的学习动态和泛化性能。具体而言,我们的主要定理表明,特征学习中梯度下降会逐步降低有效维度,将高维问题转化为一系列低维问题。这使得深度网络学习目标函数所需的样本量远低于浅层网络。尽管结果在受控训练设置下证明,我们还讨论了更常见的训练过程,并认为它们通过相同机制学习。
原文摘要 · Abstract (English)
Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of latent subspace dimensionalities. This framework enables us to analytically study the learning dynamics and generalization performance of deep networks compared to shallow ones in the high-dimensional limit. Specifically, our main theorem shows that feature learning with GD successively reduces the effective dimensionality, transforming a high-dimensional problem into a sequence of lower-dimensional ones. This enables learning the target function with drastically less samples than with shallow networks. While the results are proven in a controlled training setting, we also discuss more common training procedures and argue that they learn through the same mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。