arXiv:2409.07434stat.MLcs.LG2024-09被引 8

提出在线学习中带丢弃正则的随机梯度下降渐近理论

Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models

  • 基于几何矩收缩证明丢弃递归函数存在唯一平稳分布
  • 推导出带丢弃的ASGD与L2正则迭代差的中心极限定理
  • 设计在线估计器实现高效递归统计推断,适合大数据场景

本文为线性回归中带丢弃正则的在线随机梯度下降(SGD)迭代提出了渐近理论。具体而言,我们建立了常步长下带丢弃的SGD迭代的几何矩收缩(GMC)性质,证明了丢弃递归函数存在唯一的平稳分布。基于该性质,我们给出了丢弃迭代与ℓ²正则迭代差的逐样本中心极限定理(CLT),且不依赖初始化。同时,还给出了带丢弃的Ruppert-Polyak平均SGD(ASGD)与ℓ²正则迭代差的中心极限定理。基于这些渐近正态性结果,我们进一步提出一种在线估计器,用于递归估计ASGD丢弃的长期协方差矩阵,从而在计算时间和内存上保持高效。数值实验表明,在足够大的样本下,所提出的ASGD丢弃置信区间几乎达到名义覆盖率。

原文摘要 · Abstract (English)

This paper proposes an asymptotic theory for online inference of the stochastic gradient descent (SGD) iterates with dropout regularization in linear regression. Specifically, we establish the geometric-moment contraction (GMC) for constant step-size SGD dropout iterates to show the existence of a unique stationary distribution of the dropout recursive function. By the GMC property, we provide quenched central limit theorems (CLT) for the difference between dropout and $\ell^2$-regularized iterates, regardless of initialization. The CLT for the difference between the Ruppert-Polyak averaged SGD (ASGD) with dropout and $\ell^2$-regularized iterates is also presented. Based on these asymptotic normality results, we further introduce an online estimator for the long-run covariance matrix of ASGD dropout to facilitate inference in a recursive manner with efficiency in computational time and memory. The numerical experiments demonstrate that for sufficiently large samples, the proposed confidence intervals for ASGD with dropout nearly achieve the nominal coverage probability.

优化算法在线学习统计推断丢弃正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。