arXiv:2508.21022cs.LGmath.OC2025-08被引 1

用投影视角分析小样本自然梯度,揭示其收敛机制与优势

A Sketch-and-Project Analysis of Subsampled Natural Gradient Algorithms

  • 将小样本自然梯度视为投影视图方法,重构理论分析框架
  • 单批次即可保证全局收敛,且收敛速率与投影结构相关
  • 解释了SNG在小样本下更有效利用雅可比谱衰减的原理

子采样自然梯度下降(SNG)已被用于高精度科学机器学习,但传统基于随机预条件的分析无法解释真实的小样本情形。本文通过将SNG重新诠释为一种投影视图方法,摒弃了通常使用的双独立小批量假设,改用基于平方体积采样的新代理模型。在此框架下,我们证明即使存在梯度与预条件耦合,期望的SNG方向仍等价于预条件梯度下降步。这带来了两项关键结果:(i) 仅需任意大小的单个小批量即可实现全局收敛;(ii) 明确给出了收敛速率与投影视图结构相关的量。这些发现揭示了小样本下SNG的优势,例如能更有效地利用模型雅可比矩阵的谱衰减。此外,本文还将该思想扩展至解释一种流行的结构化动量方案SPRING,表明其可自然地从加速投影视图方法中导出。

原文摘要 · Abstract (English)

Subsampled natural gradient descent (SNG) has been used to enable high-precision scientific machine learning, but standard analyses based on stochastic preconditioning fail to provide insight into realistic small-sample settings. We overcome this limitation by instead analyzing SNG as a sketch-and-project method. Motivated by this lens, we discard the usual theoretical proxy which decouples gradients and preconditioners using two independent mini-batches, and we replace it with a new proxy based on squared volume sampling. Under this new proxy we show that the expectation of the SNG direction becomes equal to a preconditioned gradient descent step even in the presence of coupling, leading to (i) global convergence guarantees when using a single mini-batch of any size, and (ii) an explicit characterization of the convergence rate in terms of quantities related to the sketch-and-project structure. These findings in turn yield new insights into small-sample settings, for example by suggesting that the advantage of SNG over SGD is that it can more effectively exploit spectral decay in the model Jacobian. We also extend these ideas to explain a popular structured momentum scheme for SNG, known as SPRING, by showing that it arises naturally from accelerated sketch-and-project methods.

优化算法自然梯度小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。