零阶优化并非能力不足,而是被低估的高效学习方法。
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered

- 突破传统全空间设计,用子空间与谱视角降低方差
- 仅需少量查询即可实现高效训练,适合资源受限场景
- 强调前向传播优势,适用于通信高效的大规模训练
零阶(ZO)优化通过函数值有限差分进行学习,无需反向传播,近年来因内存效率高且适用于灰箱或黑箱流程而重新受到关注。尽管常被认为因估计器方差大和查询复杂度高而难以扩展,但本文认为这种判断可能源于短视的设计实践,特别是全空间、逐元素、以估计器为中心的方案。我们提出六个新视角,涵盖算法、系统与评估层面:首先,从方差控制、方差-查询权衡及方向导数角度重新审视估计器中心型方法的可行性;其次,发现三个未充分探索的机会:(i) 子空间与谱视角可实现可解释的方差缩减并支持平滑查询扩展,(ii) ZO的前向仅执行特性是系统级优势,有利于通信高效、流水线友好与资源受限训练,(iii) 需要剥离任务复杂度对ZO评估的干扰。我们主张围绕其独特优势重构ZO优化,开辟一条面向大规模、系统感知与资源高效的新型学习路径。
原文摘要 · Abstract (English)
Zeroth-order (ZO) optimization, learning from finite differences of function evaluations without backpropagation, has recently regained attention in deep learning due to its memory efficiency and applicability to gray- or black-box pipelines. Yet, ZO methods are often dismissed as fundamentally unscalable because of estimator variance and unfavorable query complexity. We argue that this conclusion might be misguided: ZO optimization is underexplored, not underpowered. We show that many perceived limitations stem from myopic development practices, most notably full-space, element-wise, estimator-centric designs. We articulate six positions spanning the algorithmic, systems, and evaluation stack. First, we revisit the feasibility boundaries of estimator-centric ZO methods through variance control, variance-query tradeoffs, and directional-derivative lenses. Then, we identify three underexplored opportunities: (i) subspace and spectral views of ZO that enable interpretable variance reduction with graceful query scaling, (ii) the forward-only nature of ZO as a systems advantage for communication-efficient, pipeline-friendly, and resource-constrained training, and (iii) the need to de-obfuscate ZO evaluations from task complexity. We strongly advocate rethinking ZO optimization around its unique strengths and acting accordingly, opening a viable path toward large-scale, system-aware, and resource-efficient learning with ZO optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。