从核视角揭示零阶优化为何能高效微调大模型
Learning Dynamics of Zeroth-Order Optimization: A Kernel Perspective

- 用经验神经正切核分析零阶梯度学习动态
- 近似误差取决于扰动次数而非参数量级
- 解释了零阶方法在大模型上的可扩展性
经典优化理论认为零阶(ZO)算法存在维数依赖的收敛瓶颈,其性能随模型维度恶化。然而,近期研究发现零阶方法可在含数十亿参数的大语言模型上成功微调。本文通过推导零阶SGD的一步学习动态,发现经验神经正切核(eNTK)是主导学习行为的关键项。分析表明,零阶eNTK中每个元素对应神经正切向量在随机低维子空间上的投影内积。结合Johnson-Lindenstrauss引理,证明其近似精度主要由扰动次数决定,且误差仅依赖于模型输出尺寸,而非参数规模。这一维数无关特性为零阶方法在大模型微调中的可扩展性提供了理论支持。该核基框架为理解零阶优化的学习动态提供了新视角。
原文摘要 · Abstract (English)
Classical optimization theory establishes that zeroth-order (ZO) algorithms suffer from a dimension-dependent slowdown, with convergence rates typically scaling with the model dimension compared to first-order methods. However, in contrast to these theoretical expectations, a growing body of recent work demonstrates the successful application of ZO methods to fine-tuning Large Language Models (LLMs) with billions of parameters. To explain this paradox, we derive the one-step learning dynamics of ZO SGD, where the empirical Neural Tangent Kernel (eNTK) naturally emerges as the key term governing the learning behavior. Inspection of the eNTK produced by ZO SGD reveals that each element corresponds to the inner product of neural tangent vectors projected onto a random low-dimensional subspace. Thus, by invoking the Johnson-Lindenstrauss Lemma, our analysis shows that the fidelity of the ZO eNTK is governed primarily by the number of perturbations. Crucially, the approximation error depends on the model output size rather than the massive parameter dimension. This dimension-free property provides a theoretical justification for the scalability of ZO methods to LLMs finetuning tasks. We believe that this kernel-based framework offers a novel perspective for understanding ZO methods within the context of learning dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。