arXiv:2409.17107math.OCcs.LG2024-09被引 1

提出可处理不连续梯度的SGHMC算法,用于训练ReLU神经网络。

Non-asymptotic convergence analysis of the stochastic gradient Hamiltonian Monte Carlo algorithm with discontinuous stochastic gradient with applications to training of ReLU neural networks

  • 允许随机梯度不连续,突破传统SGHMC限制。
  • 给出非渐近收敛上界,可控制误差至任意小。
  • 适用于金融与AI中的ReLU神经网络优化问题。

本文对随机梯度哈密顿蒙特卡洛(SGHMC)算法在Wasserstein-1和Wasserstein-2距离下收敛至目标分布提供了非渐近分析。关键在于,相比已有文献,本工作允许其随机梯度具有不连续性,从而为包含ReLU激活函数神经网络训练在内的非凸随机优化问题,提供可控制至任意小的期望过风险上界。通过数值实验,验证了方法在分位数估计及多个与金融和人工智能相关的ReLU神经网络优化问题中的适用性。

原文摘要 · Abstract (English)

In this paper, we provide a non-asymptotic analysis of the convergence of the stochastic gradient Hamiltonian Monte Carlo (SGHMC) algorithm to a target measure in Wasserstein-1 and Wasserstein-2 distance. Crucially, compared to the existing literature on SGHMC, we allow its stochastic gradient to be discontinuous. This allows us to provide explicit upper bounds, which can be controlled to be arbitrarily small, for the expected excess risk of non-convex stochastic optimization problems with discontinuous stochastic gradients, including, among others, the training of neural networks with ReLU activation function. To illustrate the applicability of our main results, we consider numerical experiments on quantile estimation and on several optimization problems involving ReLU neural networks relevant in finance and artificial intelligence.

SGHMCReLU网络非凸优化概率采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。