揭示深度随机ReLU网络在不同p范数下的最优利普希茨常数行为
Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks
- 基于随机初始化的深度神经网络,分析其ℓ^p-利普希茨常数上下界
- 浅层网络达到匹配上下界,深层网络上下界差距仅对数级于宽度
- p≥2时与高斯向量的p'范数相似,p<2时更接近ℓ²范数
本文研究了参数随机的ReLU神经网络Φ: ℝᵈ→ℝ在p∈[1,∞]下的ℓ^p-利普希茨常数。权重服从变体He初始化,偏置来自对称分布。针对宽网络,推导出高概率上下界,二者相差最多为宽度的对数因子与深度的线性因子。在浅层网络情形下获得匹配界。值得注意的是,ℓ^p-利普希茨常数在p∈[1,2)和p∈[2,∞]两个区间表现显著不同:当p≥2时,其行为类似d维标准高斯向量g的ℓ^{p'}范数(满足1/p+1/p'=1);而当p<2时,则更接近‖g‖₂。
原文摘要 · Abstract (English)
This paper studies the $\ell^p$-Lipschitz constants of ReLU neural networks $Φ: \mathbb{R}^d \to \mathbb{R}$ with random parameters for $p \in [1,\infty]$. The distribution of the weights follows a variant of the He initialization and the biases are drawn from symmetric distributions. We derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's width and linear in its depth. In the special case of shallow networks, we obtain matching bounds. Remarkably, the behavior of the $\ell^p$-Lipschitz constant varies significantly between the regimes $ p \in [1,2) $ and $ p \in [2,\infty] $. For $p \in [2,\infty]$, the $\ell^p$-Lipschitz constant behaves similarly to $\Vert g\Vert_{p'}$, where $g \in \mathbb{R}^d$ is a $d$-dimensional standard Gaussian vector and $1/p + 1/p' = 1$. In contrast, for $p \in [1,2)$, the $\ell^p$-Lipschitz constant aligns more closely to $\Vert g \Vert_{2}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。