arXiv:2506.05454cs.LGcs.AI2025-06NeurIPS被引 8

零阶优化自动找到平坦的最小值,提升模型泛化能力。

Zeroth-Order Optimization Finds Flat Minima

  • 用两点估计法实现零阶优化,隐式偏好低海森迹的解
  • 理论证明可收敛到最优解中海森迹最小的平坦最小值
  • 适用于黑箱攻击、语言模型微调等梯度难计算场景

零阶优化广泛应用于梯度不可行或计算成本高的场景,如黑箱攻击、强化学习和语言模型微调。现有理论多关注收敛至任意驻点,对最终解的隐式正则化机制了解有限。本文表明,采用标准两点估计器的零阶优化倾向于选择海森迹较小的解,该指标常用于区分尖锐与平坦最小值。进一步给出了凸函数及足够光滑函数下,零阶优化收敛至近似平坦最小值的速率。在凸损失的二分类任务和语言模型微调实验中,结果验证了理论发现。

原文摘要 · Abstract (English)

Zeroth-order methods are extensively used in machine learning applications where gradients are infeasible or expensive to compute, such as black-box attacks, reinforcement learning, and language model fine-tuning. Existing optimization theory focuses on convergence to an arbitrary stationary point, but less is known on the implicit regularization that provides a fine-grained characterization on which particular solutions are finally reached. We show that zeroth-order optimization with the standard two-point estimator favors solutions with small trace of Hessian, which is widely used in previous work to distinguish between sharp and flat minima. We further provide convergence rates of zeroth-order optimization to approximate flat minima for convex and sufficiently smooth functions, where flat minima are defined as the minimizers that achieve the smallest trace of Hessian among all optimal solutions. Experiments on binary classification tasks with convex losses and language model fine-tuning support our theoretical findings.

零阶优化平坦最小值隐式正则化黑箱攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。