arXiv:2606.00293cs.LGstat.ME2026-06被引 1

提出新方法提升大批次下模型不确定性量化精度

Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte Carlo

论文配图:Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte Carlo
图 1 · 摘自论文原文
  • 设计离散时间近似算法,改进SG(L)D采样稳定性
  • 理论证明误差有界,可准确预测协方差与自相关时间
  • 适用于大批次、模型错误设定等真实场景

在大批次或模型错误设定的现实场景中,对随机梯度下降(SGD)和随机梯度朗之万动力学(SGLD)等算法进行近似采样与不确定性量化仍具挑战。现有理论依赖连续极限或强统计假设,在这些情形下可能失准。本文提出新的带与不带动量的SG(L)D离散时间近似,可准确预测平稳协方差、迭代平均协方差及积分自相关时间。我们进一步证明了非渐近、定量的误差界,表明这些估计足以用于实际调参与不确定性量化。数值实验显示,该理论在多种模型与数据生成分布下优于现有方法,包括使用β-散度而非对数损失进行统计鲁棒推断时。

原文摘要 · Abstract (English)

Tuning algorithms such as stochastic gradient descent (SGD) and stochastic gradient Langevin dynamics (SGLD) for approximate sampling and uncertainty quantification remains challenging, particularly in the practically relevant settings when the batch size is large or the model is misspecified. Existing theory that provides tuning guidance relies on continuous-time limits or strong statistical assumptions, which can become quantitatively inaccurate in these regimes. We address these shortcomings by proposing new discrete-time approximations to SG(L)D with and without momentum, which enables accurate predictions of the stationary covariance, iterate average covariance, and integrated autocorrelation time. Moreover, we prove quantitative, non-asymptotic error bounds showing that these estimates are sufficiently accurate for practical tuning and uncertainty quantification. Numerical experiments demonstrate that our theory yields improved tuning guidance across a range of models and data-generating distributions where existing approaches fail, including when using the $β$-divergence rather than log-loss to obtain statistically robust inferences.

不确定性量化随机梯度贝叶斯推断理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。