arXiv:2503.02131stat.MLcs.LG2025-03被引 2

针对可加结构函数,提出无梯度优化新方法,但无法提升精度。

Gradient-free stochastic optimization for additive models

  • 设计随机梯度估计器,结合梯度下降实现最优误差
  • 达到 dT^(-(β-1)/β) 的最小最大优化误差
  • 适用于高维可加函数的无梯度优化问题

我们研究在噪声观测下满足Polyak-Łojasiewicz或强凸性条件的目标函数的零阶优化问题。假设目标函数具有可加结构,并满足高阶光滑性,由Hölder类函数刻画。在非参数函数估计中,已知可加模型相比一般Hölder模型能显著提高估计精度。本文将此框架应用于无梯度优化,提出一种随机梯度估计器,将其嵌入梯度下降算法后,可实现最小最大优化误差为 dT^{-(β-1)/β},其中 d 为维度,T 为查询次数,β ≥ 2 为Hölder光滑度。结论表明,在无梯度优化中,与非参数估计不同,可加结构无法带来显著精度提升。

原文摘要 · Abstract (English)

We address the problem of zero-order optimization from noisy observations for an objective function satisfying the Polyak-Łojasiewicz or the strong convexity condition. Additionally, we assume that the objective function has an additive structure and satisfies a higher-order smoothness property, characterized by the Hölder family of functions. The additive model for Hölder classes of functions is well-studied in the literature on nonparametric function estimation, where it is shown that such a model benefits from a substantial improvement of the estimation accuracy compared to the Hölder model without additive structure. We study this established framework in the context of gradient-free optimization. We propose a randomized gradient estimator that, when plugged into a gradient descent algorithm, allows one to achieve minimax optimal optimization error of the order $dT^{-(β-1)/β}$, where $d$ is the dimension of the problem, $T$ is the number of queries and $β\ge 2$ is the Hölder degree of smoothness. We conclude that, in contrast to nonparametric estimation problems, no substantial gain of accuracy can be achieved when using additive models in gradient-free optimization.

无梯度优化可加模型随机优化Hölder光滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。