arXiv:2608.06250stat.MLcs.LG2026-08

早停梯度下降可实现高斯混合分类的最小最大最优性能

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

  • 通过适时早停,梯度下降在标签噪声下仍能收敛到最优分类边界
  • 在多项式与指数谱衰减下,误差率达到理论最优水平
  • 适合研究泛化误差与优化动态关系的学者参考

在过参数化分类中,训练数据即使在分布不可分的情况下也能线性可分。此时,对逻辑损失使用梯度下降(GD)会导致范数发散,方向收敛至最大间隔插值分类器,其隐式偏差可能统计次优。本文证明:在存在标签翻转噪声的高斯混合模型中,恰当时机早停的梯度下降可实现最小最大最优的零一损失超额风险,适用于谱衰减快且连续的情形,包括多项式与指数衰减。分析结合了早停迭代的紧上界与任意分类器的匹配统计下界,实验验证了最优率。关键技术贡献是将超额逻辑损失转化为零一损失的新校准结果,处理了标签噪声引起的模型误设问题,并消除了标准界中的平方根速率。此外,我们建立了线性插值器的下界,表明插值需指数更多样本才能达到相同超额风险。

原文摘要 · Abstract (English)

In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. In this work, we show that early stopping can overcome this suboptimality: in a Gaussian mixture model with label-flipping noise, GD stopped at an appropriate oracle time achieves minimax-optimal excess zero-one risk for covariance spectra with fast and continuous decay, including polynomial and exponential spectral decays. Our analysis combines a sharp upper bound for the early-stopped iterate with a matching statistical lower bound over arbitrary classifiers, yielding optimal rates that are validated by experiments. A central technical contribution is a new calibration result that converts excess logistic risk into excess zero-one risk; it handles the model misspecification induced by the label-flipping noise, and removes the square-root rate in standard bounds. We also establish a lower bound for linear interpolators, showing that interpolation can require exponentially more samples than early stopping to achieve the same excess risk.

分类器梯度下降早停高斯混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。