分析高斯混合数据下两层神经网络一次梯度步后的表现,揭示其等价于多项式模型。
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
- 在高斯混合数据下,用渐近比例极限分析网络训练与泛化性能。
- 一次梯度步后,神经网络等价于特定阶数的多项式模型,阶数由数据分布和学习率决定。
- 实验验证了该等价性在回归、分类任务及真实数据(Fashion-MNIST)上的有效性。
本文研究在结构化数据(高斯混合模型)下,两层神经网络经过一次梯度下降后的训练与泛化性能。现有研究多基于各向同性数据假设,忽略了真实数据的复杂性。本文在输入维度、隐藏单元数与样本量以固定比例增长的渐近比例极限下,分析高斯混合数据下的两层神经网络。借助高斯普适性理论,刻画了训练与泛化误差。我们证明,在特定条件下,高阶多项式模型可等价于非线性神经网络,其等价阶数与“数据离散度”和学习率密切相关。通过大量模拟实验,验证了原模型与多项式模型在多种回归与分类任务中的等价性。此外,探讨了高斯混合的不同特性对学习效果的影响。最后,在Fashion-MNIST分类任务上展示实验结果,表明理论发现可推广至真实数据。
原文摘要 · Abstract (English)
In this work, we study the training and generalization performance of two-layer neural networks (NNs) after one gradient descent step under structured data modeled by Gaussian mixtures. While previous research has extensively analyzed this model under isotropic data assumption, such simplifications overlook the complexities inherent in real-world datasets. Our work addresses this limitation by analyzing two-layer NNs under Gaussian mixture data assumption in the asymptotically proportional limit, where the input dimension, number of hidden neurons, and sample size grow with finite ratios. We characterize the training and generalization errors by leveraging recent advancements in Gaussian universality. Specifically, we prove that a high-order polynomial model performs equivalent to the nonlinear neural networks under certain conditions. The degree of the equivalent model is intricately linked to both the "data spread" and the learning rate employed during one gradient step. Through extensive simulations, we demonstrate the equivalence between the original model and its polynomial counterpart across various regression and classification tasks. Additionally, we explore how different properties of Gaussian mixtures affect learning outcomes. Finally, we illustrate experimental results on Fashion-MNIST classification, indicating that our findings can translate to realistic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。