揭示高维经验风险最小化中高斯普适性失效的边界条件
Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization
- 扩展凸高斯极小极大定理至非高斯数据,推导出估计量统计特性的渐近极小极大刻画
- 证明测试投影服从非高斯分布与独立高斯噪声的卷积,其方差由协方差矩阵决定
- 适用于分析非高斯设计下各类损失函数和正则化的高维统计行为
研究在一般非高斯数据设计下的高维凸经验风险最小化(ERM)。通过启发式地将凸高斯极小极大定理(CGMT)推广到非高斯情形,我们推导出关键统计量的渐近极小极大刻画,可近似估计ERM估计量 $\hatθ$ 的均值 $μ_{\hatθ}$ 和协方差 $C_{\hatθ}$。在数据矩阵满足集中性假设且损失函数和正则项满足标准正则性条件下,对于与训练数据独立的测试协变量 $x$,投影 $\hatθ^\top x$ 近似服从 $μ_{\hatθ}^\top x$ 的非高斯分布与一个独立中心高斯变量(方差为 $\mathrm{tr}(C_{\hatθ} \mathbb{E}[xx^\top])$)的卷积。该结果澄清了高斯普适性在ERM中的适用范围与局限。通过多种损失函数和模型的数值模拟验证了理论预测与定性洞察。
原文摘要 · Abstract (English)
We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min-max characterization of key statistics, enabling approximation of the mean $μ_{\hatθ}$ and covariance $C_{\hatθ}$ of the ERM estimator $\hatθ$. Specifically, under a concentration assumption on the data matrix and standard regularity conditions on the loss and regularizer, we show that for a test covariate $x$ independent of the training data, the projection $\hatθ^\top x$ approximately follows the convolution of the generally non-Gaussian distribution of $μ_{\hatθ}^\top x$ with an independent centered Gaussian variable of variance $\mathrm{tr}(C_{\hatθ} \mathbb{E}[xx^\top])$. This result clarifies the scope and limits of Gaussian universality for ERMs. Numerical simulations across diverse losses and models are provided to validate our theoretical predictions and qualitative insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。