arXiv:2510.14074stat.MLcs.LG2025-10被引 1

该论文精确解析了多类数据下SGD的动态过程,揭示了数据异质性对学习效果的关键影响。

Exact Dynamics of Multi-class Stochastic Gradient Descent

  • 通过高维极限下的常微分方程组,精确描述了多类SGD的学习轨迹。
  • 在最小二乘和二分类逻辑回归中,发现学习率存在依赖于协方差均值特征的阈值。
  • 揭示数据异质性引发结构相变,使模型更关注低方差方向的类别均值,适用于高维建模研究者。

我们构建了一个分析高维问题中单遍随机梯度下降(SGD)学习动态的框架,数据来自多个各向异性类别。主定理在高维极限下,以确定性常微分方程组形式给出了风险与真实信号重叠等关键量的精确表达式,适用于广泛优化问题,并可扩展至类别数随维度增长的情形。为展示其效用,我们详细研究了数据各向异性在二分类逻辑回归和最小二乘损失中的影响。在线性多类设定下,推导出依赖于协方差矩阵平均特征值的学习率阈值。在二分类逻辑回归中,考察三种情形:各向同性协方差、具有大量零特征值的协方差矩阵(零一模型)、以及幂律谱协方差矩阵。结果显示存在结构性相变:在零一模型和足够大幂律指数的幂律模型中,SGD更倾向于对低方差方向(即“干净方向”)上的类别均值进行投影。该结论经解析分析与数值模拟验证,准确反映了高维极限下的损失渐近行为。这些数据异质性效应可能普遍存在于更广泛场景,体现了所证定理的普适应用价值。

原文摘要 · Abstract (English)

We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes. Our main theorem provides exact expressions for quantities of interest, including the risk and the overlap with the true signal, in terms of a deterministic system of ODEs, valid in the high-dimensional limit. The theorem holds for a broad class of optimization problems and extends to settings where the number of classes grows with dimension. To illustrate its utility, we investigate in detail the effect of the data's anisotropic structure on the problems of binary logistic regression and least-squares (LS) loss. We study the LS in a linear multiclass setup and derive a learning-rate threshold that depends on the average eigenvalue of the covariance matrices. In the binary logistic regression, we study three cases: isotropic covariances, data covariance matrices with a large fraction of zero eigenvalues (denoted as the zero-one model), and covariance matrices with power-law spectra. We show that a structural phase transition occurs. In particular, for the zero-one model and the power-law model with sufficiently large power, SGD aligns more closely with values of the class mean that are projected onto the ``clean directions'' (i.e., directions of smaller variance). This is supported by analytical studies and numerical simulations, which show the exact asymptotic behavior of the loss in the high-dimensional limit. The effects of data anisotropy that we demonstrate are likely to hold beyond these examples and illustrate one application of the broader theorem that we prove.

机器学习优化理论高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。