arXiv:2608.03368cs.LG2026-08

揭示ReLU神经网络核矩阵最小特征值的紧致下界及其最优性。

Tight Worst-Case Bounds for the Smallest Eigenvalue of ReLU NTK Gram Matrices

  • 基于向量间投影距离,建立无维度依赖的最小特征值下界。
  • 证明该下界在最坏情况下可被逼近,达到理论最优。
  • 对理解深层网络泛化能力有重要意义,适合理论研究者。

对于 $n$ 个单位向量 $x_1,\ dots,x_n \in \mathbb{R}^d$,我们研究连续ReLU导数核矩阵 $H$,其元素为沿标准高斯方向平均的成对门控内积。定义 $ Δ_\pm := \min_{i \neq j} \min\{ \|x_i-x_j\|_2, \|x_i+x_j\|_2 \} $ 为其投影间距,我们证明了普适的无维度下界 $ λ_{\min}(H) = Ω( Δ_\pm/\sqrt{\log n} ) $。同时构造出满足匹配上界 $ λ_{\min}(H) = O( Δ_\pm/\sqrt{\log n} ) $ 的最坏情况族,表明该速率在常数因子意义下是紧致的。

原文摘要 · Abstract (English)

For $n$ unit vectors $x_1,\ldots,x_n \in \mathbb{R}^d$, we study the continuous ReLU derivative Gram matrix $H$, whose entries are obtained by averaging pairwise gated inner products over a standard Gaussian direction. Writing $ Δ_\pm := \min_{i \neq j} \min\{ \|x_i-x_j\|_2, \|x_i+x_j\|_2 \} $ for their projective separation, we prove the universal dimension-free lower bound $ λ_{\min}(H) = Ω( Δ_\pm/\sqrt{\log n} ) $. Conversely, we construct worst-case families satisfying the matching upper bound $ λ_{\min}(H) = O( Δ_\pm/\sqrt{\log n} ) $, showing that this rate is tight up to universal constants.

神经网络谱分析特征值理论深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。