arXiv:2606.06469math.STcs.LG2026-06被引 1

绝大多数能正确分类的模型,泛化性能高度一致。

How abundant are good interpolators?

  • 在高维数据下,所有正确分类的线性模型中,性能相近
  • 超过99%的分类器泛化误差集中在极小范围
  • 适合研究过参数模型的泛化机制

设S为所有单位范数的线性分类器θ∈ℝᵈ,它们能正确分类给定数据集{(Xᵢ,yᵢ)}ₙᵢ₌₁(Xᵢ∈ℝᵈ,yᵢ∈{−1,+1}),且最小边距κ预先固定。在高斯混合与逻辑回归(高斯特征)两类数据生成模型下,在n/d→α且α较小时,我们建立了随机选取θ∈S时其泛化误差的典型大偏差原理。该偏差率函数为确定性函数,描述了在指数尺度上具有特定性能的插值分类器的比例。结果表明:几乎所有插值分类器(除指数级小部分外)的泛化性能几乎相同,由该率函数唯一最大值决定。数值实验显示,梯度下降和线性规划所找到的解虽在集合S中,但性能显著优于绝大多数插值分类器,说明在过参数化情形下这些方法具有非平凡的良性过拟合现象。

原文摘要 · Abstract (English)

Let $S$ be the set of unit norm linear classifiers $θ\in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $κ$ fixed in advance. Under two natural data-generating distributions of the $(X,y)$ pairs -- a Gaussian mixture model and a logistic model with Gaussian features -- and in the proportional regime $n/d \to α$ with small enough $α$, we establish a large deviation principle on the event that a point $θ$ chosen uniformly at random from $S$ achieves a given generalization error, with high probability over the choice of the data. The associated large deviation rate function is deterministic and describes the proportion, at the exponential scale in $d$, of interpolating classifiers having a given desired performance. As a consequence, we establish the following concentration phenomenon: all but an exponentially small fraction of interpolating classifiers have approximately the same generalization performance given by the unique maximizer of this rate function. We numerically compare this maximizer to the performance of empirical risk minimization by gradient descent and to the performance of a natural linear program, both finding a point in $S$, and deduce that in the overparametrized regime of small $α$, these efficient procedures outperform the vast majority of interpolators, pointing to their nontrivial benign overfitting in this setting.

泛化分析过参数化插值分类器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。