arXiv:2603.21169cs.LG2026-03

提出新型核方法解析零阶优化训练动态,揭示其收敛机制。

Model Evolution Under Zeroth-Order Optimization: A Neural Tangent Kernel Perspective

  • 引入神经零阶核(NZK)刻画模型在函数空间的演化轨迹。
  • 证明线性模型下预期NZK恒定,可导出平方损失下的闭式解。
  • 实验验证理论并发现单个随机向量能加速收敛,适合低内存场景。

零阶(ZO)优化通过仅依赖前向传播估计梯度,无需反向传播即可实现内存高效训练神经网络。然而,梯度估计的随机性严重模糊了训练动态,与一阶方法在神经正切核(NTK)理论下的清晰行为形成对比。为此,我们引入神经零阶核(NZK)以描述在ZO更新下模型在函数空间中的演化。对于线性模型,我们证明预期NZK在整个训练过程中保持不变,且明确依赖于随机扰动方向的一阶和二阶矩。该不变性导出了平方损失下模型演化的闭式表达式。进一步将分析扩展至线性化神经网络。将ZO更新解释为基于NZK的核梯度下降,为可能加速收敛提供了新视角。在合成数据及真实世界数据集(包括MNIST、CIFAR-10和Tiny ImageNet)上的大量实验验证了理论结果,并表明使用单一共享随机向量可实现加速。

原文摘要 · Abstract (English)

Zeroth-order (ZO) optimization enables memory-efficient training of neural networks by estimating gradients via forward passes only, eliminating the need for backpropagation. However, the stochastic nature of gradient estimation significantly obscures the training dynamics, in contrast to the well-characterized behavior of first-order methods under Neural Tangent Kernel (NTK) theory. To address this, we introduce the Neural Zeroth-order Kernel (NZK) to describe model evolution in function space under ZO updates. For linear models, we prove that the expected NZK remains constant throughout training and depends explicitly on the first and second moments of the random perturbation directions. This invariance yields a closed-form expression for model evolution under squared loss. We further extend the analysis to linearized neural networks. Interpreting ZO updates as kernel gradient descent via NZK provides a novel perspective for potentially accelerating convergence. Extensive experiments across synthetic and real-world datasets (including MNIST, CIFAR-10, and Tiny ImageNet) validate our theoretical results and demonstrate acceleration when using a single shared random vector.

零阶优化核方法训练动态内存效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。