arXiv:2512.03325math.STcs.LG2025-12被引 6

揭示随机特征模型在二次缩放下高斯等价失效机制并提出修正方案

When does Gaussian equivalence fail and how to fix it: Non-universal behavior of random features with quadratic scaling

  • 提出条件高斯等价模型,融合低维非高斯成分修正传统高斯近似
  • 在数据维度与样本量平方增长时,准确预测训练/测试误差的渐近行为
  • 适用于低维投影目标函数场景,为高维经验风险最小化提供新分析工具

现代高维统计中,线性预测器通过经验风险最小化(ERM)在非线性特征嵌入上训练的研究备受关注。高斯等价理论(GET)作为强大的普适性原理,认为复杂高维特征的行为可由更易分析的高斯替代模型捕捉。然而,数值实验显示,即使在简单嵌入(如多项式映射)下,通用缩放条件下该等价性仍可能失效。本文研究随机特征(RF)模型在二次缩放情形下的这一失效问题,即当特征数与样本量均随数据维度平方增长时。我们发现,若目标函数依赖于数据的低维投影(如广义线性模型),标准GET会产生错误预测。为此,我们引入条件高斯等价(CGE)模型:在高维高斯框架中附加一个低维非高斯分量。该混合模型保持了高斯框架的可处理性,并能精确描述二次缩放下的RF模型。我们推导出训练与测试误差的尖锐渐近表达式,其结果在标准GET失效时仍与数值模拟高度吻合。分析结合了关于维纳混沌展开的中心极限定理及精细的两阶段Lindeberg交换论证。本工作不仅拓展至随机特征模型与二次缩放,更揭示了高维ERM中丰富的普适性现象图景。

原文摘要 · Abstract (English)

A major effort in modern high-dimensional statistics has been devoted to the analysis of linear predictors trained on nonlinear feature embeddings via empirical risk minimization (ERM). Gaussian equivalence theory (GET) has emerged as a powerful universality principle in this context: it states that the behavior of high-dimensional, complex features can be captured by Gaussian surrogates, which are more amenable to analysis. Despite its remarkable successes, numerical experiments show that this equivalence can fail even for simple embeddings -- such as polynomial maps -- under general scaling regimes. We investigate this breakdown in the setting of random feature (RF) models in the quadratic scaling regime, where both the number of features and the sample size grow quadratically with the data dimension. We show that when the target function depends on a low-dimensional projection of the data, such as generalized linear models, GET yields incorrect predictions. To capture the correct asymptotics, we introduce a Conditional Gaussian Equivalent (CGE) model, which can be viewed as appending a low-dimensional non-Gaussian component to an otherwise high-dimensional Gaussian model. This hybrid model retains the tractability of the Gaussian framework and accurately describes RF models in the quadratic scaling regime. We derive sharp asymptotics for the training and test errors in this setting, which continue to agree with numerical simulations even when GET fails. Our analysis combines general results on CLT for Wiener chaos expansions and a careful two-phase Lindeberg swapping argument. Beyond RF models and quadratic scaling, our work hints at a rich landscape of universality phenomena in high-dimensional ERM.

随机特征高维统计普适性误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。