arXiv:2410.02629math.STcs.LG2024-10NeurIPS

提出可追踪迭代过程泛化误差的估计方法,适用于高维鲁棒回归。

Estimating Generalization Performance Along the Trajectory of Proximal SGD in Robust Regression

  • 基于近端SGD设计泛化误差跟踪器,支持非光滑正则项。
  • 在重尾误差下,对梯度下降和随机梯度下降的迭代点给出精确误差估计。
  • 可指导最优停止迭代,适合高维鲁棒回归研究者使用。

本文研究在高维鲁棒回归问题中,梯度下降(GD)、随机梯度下降(SGD)及其近端变体所产生迭代点的泛化性能。特征数与样本量相当,且误差可能服从重尾分布。本文提出可精确追踪算法迭代轨迹上泛化误差的估计器,在适当条件下具有理论一致性。通过多个例子验证,包括Huber回归、伪Huber回归及其带非光滑正则化的变体。给出了来自GD、SGD或含非光滑正则项的近端SGD生成迭代点的显式泛化误差估计。所提出的风险估计能有效逼近真实泛化误差,从而确定最小化泛化误差的最优停止迭代。大量模拟实验验证了该估计方法的有效性。

原文摘要 · Abstract (English)

This paper studies the generalization performance of iterates obtained by Gradient Descent (GD), Stochastic Gradient Descent (SGD) and their proximal variants in high-dimensional robust regression problems. The number of features is comparable to the sample size and errors may be heavy-tailed. We introduce estimators that precisely track the generalization error of the iterates along the trajectory of the iterative algorithm. These estimators are provably consistent under suitable conditions. The results are illustrated through several examples, including Huber regression, pseudo-Huber regression, and their penalized variants with non-smooth regularizer. We provide explicit generalization error estimates for iterates generated from GD and SGD, or from proximal SGD in the presence of a non-smooth regularizer. The proposed risk estimates serve as effective proxies for the actual generalization error, allowing us to determine the optimal stopping iteration that minimizes the generalization error. Extensive simulations confirm the effectiveness of the proposed generalization error estimates.

鲁棒回归泛化误差近端SGD高维统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。