研究高维下二次神经网络的优化与样本复杂度,揭示训练动态与泛化性能的量化规律。
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
- 通过分析矩阵Riccati微分方程,刻画梯度下降的特征学习过程。
- 发现预测风险随时间、样本量和模型宽度呈幂律变化,且可精确计算。
- 适用于理解高维深度学习中优化行为的理论研究者或算法设计者。
我们研究在高维情形下,基于梯度的两层二次激活神经网络的优化与样本复杂度问题。数据生成方式为 $f_*(oldsymbol{x}) /propto extstyleinom{r}{j=1} λ_j σ(⟨oldsymbol{θ_j}, oldsymbol{x}⟩)$,其中 $oldsymbol{x} ilde{N}(0,oldsymbol{I}_d)$,$σ$ 为二阶埃尔米特多项式,$oldsymbol{θ_j}$ 为正交信号方向。考虑广义宽度情形 $r acksim d^β$($βackslashin [0,1)$),并假设第二层权重满足幂律衰减 $λ_j acksim j^{-α}$($αackslashgeq 0$)。本文对种群极限与有限样本(在线)离散化下的随机梯度下降动态进行精确分析,推导出预测风险的尺度律,揭示其对优化时间、样本量与模型宽度的幂律依赖关系。分析结合了矩阵Riccati微分方程的精细刻画与新颖的矩阵单调性论证,建立了无限维有效动态的收敛保证。
原文摘要 · Abstract (English)
We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as $f_*(\boldsymbol{x}) \propto \sum_{j=1}^{r}λ_j σ\left(\langle \boldsymbol{θ_j}, \boldsymbol{x}\rangle\right), \boldsymbol{x} \sim N(0,\boldsymbol{I}_d)$, $σ$ is the 2nd Hermite polynomial, and $\lbrace\boldsymbolθ_j \rbrace_{j=1}^{r} \subset \mathbb{R}^d$ are orthonormal signal directions. We consider the extensive-width regime $r \asymp d^β$ for $β\in [0, 1)$, and assume a power-law decay on the (non-negative) second-layer coefficients $λ_j\asymp j^{-α}$ for $α\geq 0$. We present a sharp analysis of the SGD dynamics in the feature learning regime, for both the population limit and the finite-sample (online) discretization, and derive scaling laws for the prediction risk that highlight the power-law dependencies on the optimization time, sample size, and model width. Our analysis combines a precise characterization of the associated matrix Riccati differential equation with novel matrix monotonicity arguments to establish convergence guarantees for the infinite-dimensional effective dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。