arXiv:2410.05609stat.MLcs.LG2024-10ICLR被引 6

揭示高维混合数据下高斯普适性失效现象,指出分类性能受分布细节影响。

The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures

  • 基于线性因子混合模型,精确刻画高维分类中经验风险最小化行为
  • 发现高斯普适性在该设置下失效,性能依赖于超出均值与协方差的分布信息
  • 明确普适性成立条件,指导损失函数选择,适合研究高维统计学习理论者

高维机器学习中,高斯或高斯混合数据假设被广泛用于精确分析方法性能。为放宽此限制,研究提出“高斯等价原理”,即在某些情况下,非高斯数据的渐近性能与同均值、协方差的高斯数据一致。然而,高斯普适性之外,关于数据分布如何影响学习性能的精确结果仍极少。本文针对一类扩展高斯混合的线性因子混合数据,对分类任务中的经验风险最小化进行高维精确刻画。结果表明,高斯普适性在此设定下失效:渐近分类性能不仅取决于类别均值和协方差,还受更深层分布结构影响。文章进一步给出高斯普适性成立的条件,并讨论其对损失函数选择的启示。

原文摘要 · Abstract (English)

The assumption of Gaussian or Gaussian mixture data has been extensively exploited in a long series of precise performance analyses of machine learning (ML) methods, on large datasets having comparably numerous samples and features. To relax this restrictive assumption, subsequent efforts have been devoted to establish "Gaussian equivalent principles" by studying scenarios of Gaussian universality where the asymptotic performance of ML methods on non-Gaussian data remains unchanged when replaced with Gaussian data having the same mean and covariance. Beyond the realm of Gaussian universality, there are few exact results on how the data distribution affects the learning performance. In this article, we provide a precise high-dimensional characterization of empirical risk minimization, for classification under a general mixture data setting of linear factor models that extends Gaussian mixtures. The Gaussian universality is shown to break down under this setting, in the sense that the asymptotic learning performance depends on the data distribution beyond the class means and covariances. To clarify the limitations of Gaussian universality in the classification of mixture data and to understand the impact of its breakdown, we specify conditions for Gaussian universality and discuss their implications for the choice of loss function.

高维统计分类性能高斯普适性线性因子模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。