arXiv:2508.20326stat.MLcs.LG2025-08NeurIPS被引 2

在存在干扰参数时,仍能保证随机梯度算法收敛。

Stochastic Gradients under Nuisances

论文配图:Stochastic Gradients under Nuisances
图 1 · 摘自论文原文
  • 引入干扰参数下的随机梯度优化框架,分析其收敛性。
  • 证明在奈曼正交条件下,经典算法仍可收敛。
  • 提出近似正交化更新变体,适用于非正交场景。

随机梯度优化是多种学习场景(从经典监督学习到现代自监督学习)的主流范式。本文研究目标依赖于未知干扰参数的学习问题中的随机梯度算法,并建立了非渐近收敛保证。结果表明,尽管干扰参数可能改变最优解并扰乱优化轨迹,但在满足适当条件(如奈曼正交性)时,经典随机梯度算法仍可收敛。即使不满足奈曼正交性,我们还证明一种具有近似正交化更新的算法变体(使用近似正交化梯度算子)也能达到类似的收敛速率。文中讨论了正交统计学习/双重机器学习和因果推断等实例。

原文摘要 · Abstract (English)

Stochastic gradient optimization is the dominant learning paradigm for a variety of scenarios, from classical supervised learning to modern self-supervised learning. We consider stochastic gradient algorithms for learning problems whose objectives rely on unknown nuisance parameters, and establish non-asymptotic convergence guarantees. Our results show that, while the presence of a nuisance can alter the optimum and upset the optimization trajectory, the classical stochastic gradient algorithm may still converge under appropriate conditions, such as Neyman orthogonality. Moreover, even when Neyman orthogonality is not satisfied, we show that an algorithm variant with approximately orthogonalized updates (with an approximately orthogonalized gradient oracle) may achieve similar convergence rates. Examples from orthogonal statistical learning/double machine learning and causal inference are discussed.

优化理论因果推断随机梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。