arXiv:2603.14578stat.MLcs.LG2026-03被引 2

揭示随机特征模型中数据谱结构如何保留幂律特性

Power-Law Spectrum of the Random Feature Model

  • 分析随机线性投影加非线性激活后数据谱的演化规律
  • 证明输入的幂律谱在输出中被精确保留,仅受对数修正
  • 适用于研究神经网络缩放定律与谱结构关系的研究者

神经网络的缩放定律中,损失随参数量、数据量和计算量呈幂律下降,其关键在于数据协方差的谱结构。视觉与语言任务中普遍出现幂律衰减的特征值。核心问题是:当数据通过神经网络的基本单元——随机线性投影加非线性激活时,这种谱结构是否被破坏?本文研究随机特征模型:设输入 $x \sim N(0,H)\in \mathbb{R}^v$,其中协方差矩阵 $H$ 具有 $\alpha$-幂律谱($\lambda_j(H) \asymp j^{-\alpha}$,$\alpha>1$),使用高斯稀疏矩阵 $W \in \mathbb{R}^{v\times d}$,以及逐元素幂函数 $f(y) = y^{p}$,我们刻画了群体随机特征协方差 $\mathbb{E}_{x }[\frac{1}{d}f(W^\top x )^{\otimes 2}]$ 的特征值。证明了上下界匹配:对所有 $1 \leq j \leq c_1 d \log^{-(p+1)}(d)$,第 $j$ 个特征值为 $\left(\log^{p-1}(j+1)/j\right)^\alpha$;对 $c_1 d \log^{-(p+1)}(d)\leq j\leq d$,第 $j$ 个特征值为 $j^{-\alpha}$,至多一个多项式对数因子。即输入的幂律指数 $\alpha$ 被精确继承,仅受单次幂次 $p$ 决定的对数修正。证明结合了分段头尾分解、高阶沃尔克混沌展开及随机矩阵集中不等式。

原文摘要 · Abstract (English)

Scaling laws for neural networks, in which the loss decays as a power-law in the number of parameters, data, and compute, depend fundamentally on the spectral structure of the data covariance, with power-law eigenvalue decay appearing ubiquitously in vision and language tasks. A central question is whether this spectral structure is preserved or destroyed when data passes through the basic building block of a neural network: a random linear projection followed by a nonlinear activation. We study this question for the random feature model: given data $x \sim N(0,H)\in \mathbb{R}^v$ where $H$ has $α$-power-law spectrum ($λ_j(H ) \asymp j^{-α}$, $α> 1$), a Gaussian sketch matrix $W \in \mathbb{R}^{v\times d}$, and an entrywise monomial $f(y) = y^{p}$, we characterize the eigenvalues of the population random-feature covariance $\mathbb{E}_{x }[\frac{1}{d}f(W^\top x )^{\otimes 2}]$. We prove matching upper and lower bounds: for all $1 \leq j \leq c_1 d \log^{-(p+1)}(d)$, the $j$-th eigenvalue is of order $\left(\log^{p-1}(j+1)/j\right)^α$. For $ c_1 d \log^{-(p+1)}(d)\leq j\leq d$, the $j$-th eigenvalue is of order $j^{-α}$ up to a polylog factor. That is, the power-law exponent $α$ is inherited exactly from the input covariance, modified only by a logarithmic correction that depends on the monomial degree $p$. The proof combines a dyadic head-tail decomposition with Wick chaos expansions for higher-order monomials and random matrix concentration inequalities.

随机特征幂律谱神经网络缩放随机矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。