解决梯度方差无穷时的优化不确定性问题,给出可靠置信区间。
Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance
- 用Polyak-Ruppert平均与自标准化统计量,突破传统方差有限假设限制。
- 在方差有限和无限两种情况下均实现渐近有效置信区域。
- 无需估计尾指数等复杂参数,适合实际优化中的不确定性量化。
随机梯度下降(SGD)是大规模统计学习和随机优化的基础。但在某些现代学习问题中,随机梯度可能呈现无穷方差行为,导致依赖有限方差假设的经典推断方法失效。本文提出一种模型无关的方法,可在有限与无限方差情形下构建基于SGD迭代的置信区域。首先证明,Polyak-Ruppert平均的渐近方向尺度不超过最快收敛的最终迭代点,类似于有限方差情况下的更低渐近方差。因此,推断方法聚焦于Polyak-Ruppert平均估计器。我们建立了该估计器与同一迭代序列中经验二阶矩归一化器的联合中心极限定理,由此生成自标准化统计量,其中主导的尾部相关缩放项相互抵消。随后采用子采样估计相关分位数,避免了对异构参数(如尾指数、慢变函数或稳定分布参数)的显式估计。所得置信区域实现简单且在两种情形下渐近有效。实证研究显示在多种设置下具有可靠覆盖性,支持该方法作为随机优化中不确定性量化的一种实用工具。
原文摘要 · Abstract (English)
Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization. However, in some modern statistical learning problems, stochastic gradients can exhibit infinite-variance behavior. Consequently, classical inference methods for SGD that rely on a finite-variance assumption break down. We develop a model-agnostic methodology for constructing confidence regions from SGD iterates in both the finite- and infinite-variance regimes. We first show that Polyak--Ruppert averaging has an asymptotic directional scale no larger than that of the fastest-rate final iterate, analogous to its lower asymptotic variance in the finite-variance setting. Accordingly, we focus our inference methodology on the Polyak--Ruppert averaged estimator. Specifically, we establish a joint central limit theorem for this estimator and an empirical second-moment normalizer from the same iterates. This joint limit yields a self-normalized statistic in which the leading tail-dependent scaling terms cancel. We then use subsampling to estimate the relevant quantiles, avoiding explicit estimation of nuisance parameters including tail indices, slowly varying functions, or stable-law parameters. The resulting confidence regions are straightforward to implement and asymptotically valid in both the finite- and infinite-variance regimes. Empirical studies show reliable coverage in various settings, supporting the proposed method as a practical tool for uncertainty quantification in stochastic optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。