arXiv:2502.09106cs.LG2025-02被引 1

研究线性回归中随机梯度下降的缩放规律,揭示特征学习对泛化性能的影响。

Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression

  • 采用二次参数化线性回归模型,分析SGD在高维数据下的收敛特性。
  • 发现模型变量的学习率会自动匹配真值衰减规律,实现自适应拟合。
  • 首次在理论上区分有无特征学习的泛化曲线,适用于理解深层网络训练机制。

在机器学习中,缩放定律描述了模型性能随模型规模和数据量增大而提升的规律。从学习理论视角看,这类结果为特定学习算法建立了泛化上界与下界。在此类研究中,具体算法与特定参数化方式常带来关键的隐式正则化效应,从而实现良好泛化。以往理论工作主要聚焦于线性模型,而神经网络取得显著经验成功的关键过程——特征学习,却长期缺失理论刻画。本文研究一种二次参数化的线性回归模型中的缩放定律,考虑无穷维数据与具有幂律衰减特性的真值斜率。我们分析了随机梯度下降(SGD)的收敛速率,并证明变量学习率会自动适配真值衰减速率。因此,在经典线性回归框架下,我们明确区分了有无特征学习时的泛化曲线差异,并给出了不依赖参数化方法和算法的信息论下界。对衰减真值的分析,为模型学习动态提供了新的刻画视角。

原文摘要 · Abstract (English)

In machine learning, the scaling law describes how the model performance improves with the model and data size scaling up. From a learning theory perspective, this class of results establishes upper and lower generalization bounds for a specific learning algorithm. Here, the exact algorithm running using a specific model parameterization often offers a crucial implicit regularization effect, leading to good generalization. To characterize the scaling law, previous theoretical studies mainly focus on linear models, whereas, feature learning, a notable process that contributes to the remarkable empirical success of neural networks, is regretfully vacant. This paper studies the scaling law over a linear regression with the model being quadratically parameterized. We consider infinitely dimensional data and slope ground truth, both signals exhibiting certain power-law decay rates. We study convergence rates for Stochastic Gradient Descent and demonstrate the learning rates for variables will automatically adapt to the ground truth. As a result, in the canonical linear regression, we provide explicit separations for generalization curves between SGD with and without feature learning, and the information-theoretical lower bound that is agnostic to parametrization method and the algorithm. Our analysis for decaying ground truth provides a new characterization for the learning dynamic of the model.

缩放定律随机梯度下降线性回归特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。