研究随机权重对线性回归优化的影响,揭示噪声与统计性能的权衡。
The Interplay of Statistics and Noisy Optimization: Learning Linear Predictors with Random Data Weights
- 用随机数据权重统一分析多种优化噪声,建立理论框架。
- 证明了收敛速度在均值和方差上的非渐近界,明确噪声影响。
- 发现快收敛的权重可能导致统计性能下降,提醒实践者注意平衡。
我们在一个通用权重分布下分析了带随机加权数据点的梯度下降在线性回归模型中的表现。该框架涵盖多种随机梯度下降、重要性采样,以及任意连续取值的权重分布,统一研究各类噪声对训练轨迹的影响。我们刻画了随机加权带来的隐式正则化,将其与加权线性回归关联,并推导出一阶与二阶矩收敛的非渐近界。利用几何矩收缩方法,进一步研究了噪声引入的平稳分布。基于这些结果,讨论了特定权重分布如何影响优化问题本质与估计器的统计性质,揭示了某些加速收敛的权重选择反而导致较差的统计性能的实例。
原文摘要 · Abstract (English)
We analyze gradient descent with randomly weighted data points in a linear regression model, under a generic weighting distribution. This includes various forms of stochastic gradient descent, importance sampling, but also extends to weighting distributions with arbitrary continuous values, thereby providing a unified framework to analyze the impact of various kinds of noise on the training trajectory. We characterize the implicit regularization induced through the random weighting, connect it with weighted linear regression, and derive non-asymptotic bounds for convergence in first and second moments. Leveraging geometric moment contraction, we also investigate the stationary distribution induced by the added noise. Based on these results, we discuss how specific choices of weighting distribution influence both the underlying optimization problem and statistical properties of the resulting estimator, as well as some examples for which weightings that lead to fast convergence cause bad statistical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。