arXiv:2608.26526cs.LGcs.NA2026-08

为宽随机tanh网络建立了高概率导数界,揭示了深度无关的平滑性。

High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube

  • 通过控制梯度方向的有限覆盖,抑制了深层网络导数的指数增长
  • 一阶导数与深度无关,高阶混合导数以多项式速度增长
  • 结果可直接用于基于拟蒙特卡罗的训练分析,适合理论研究者

本文建立了满足阶乘增长条件的激活函数下,宽随机神经网络混合输入导数的高概率界。针对采用Xavier初始化的tanh网络,证明当隐藏层宽度n满足n ≥ C(L³n₀²(1+log n₀)+L²(1+log(L/η)))时,以至少1−η的概率,对任意非空u⊆[n₀]和x∈[0,1]ⁿ₀,均有|Dᵘℛ_{Φ⁽ᴸ⁾}(x)| ≤ C₀|u|!(C₁L)^(|u|−1)∏_{j∈u}β_j(η,n₀)。这意味着一阶导数与深度无关,|u|阶混合导数至多以L^{|u|−1}多项式增长(忽略坐标因子)。由此推导出网络实现的欧氏利普希茨常数与加权索博列夫范数的高概率界,这些结果将导数估计与拟蒙特卡罗积分联系起来,指明了此类正则性在基于QMC的训练分析中的应用路径。

原文摘要 · Abstract (English)

We establish high-probability bounds for mixed input derivatives of wide random neural networks whose activation derivatives satisfy a factorial growth bound. Our main result specializes these estimates to $\tanh$ networks with Xavier initialization. A direct deterministic analysis based on Euclidean operator norms of the weight matrices yields derivative bounds that generally grow exponentially with the depth. We show that this growth can be substantially improved for sufficiently wide Gaussian networks by isolating the term that is linear in the highest-order derivative and controlling the corresponding tangent directions by measurable finite nets. For scalar-output $\tanh$ networks with Gaussian weights and Xavier initialization, we prove that there exist constants $C,C_0,C_1>0$ such that, whenever the common hidden width satisfies $n \geq C\left(L^3n_0^2(1+\log n_0)+L^2\left(1+\log(L/η)\right)\right)$, then, with probability at least $1-η$, the estimate $\left|D^u\mathcal{R}_{Φ^{(L)}}(x)\right| \leq C_0 |u|! (C_1L)^{|u|-1}\prod_{j\in u}β_j(η,n_0)$ holds simultaneously for every non-empty $u\subseteq[n_0]$ and every $x\in[0,1]^{n_0}$. Thus, the first-order derivative bound is independent of the depth, while a square-free mixed derivative of order $|u|$ grows at most polynomially as $L^{|u|-1}$, apart from the coordinate factors. As consequences, we obtain high-probability bounds for the Euclidean Lipschitz constant and for weighted Sobolev norms of the network realization. The latter connect the derivative estimates to quasi-Monte Carlo integration and indicate how such regularity can enter the analysis of QMC-based training.

神经网络导数界随机初始化QMC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。