arXiv:2506.04878math.STcs.LG2025-06被引 3

提出kTULA算法,解决深度学习中对数梯度超线性增长的采样难题

kTULA: A Langevin sampling algorithm with improved KL bounds under super-linear log-gradients

  • 基于修正的Langevin动力学,通过引入衰减机制应对超线性梯度
  • 在KL散度上实现2−ε的最优收敛速率,优于已有方法
  • 适用于高维双阱势和神经网络优化,可提供理论性能保证

针对深度学习中全局Lipschitz连续性不成立的情况,研究对数梯度超线性增长分布的采样问题。提出一种基于驯化Langevin动力学的新算法kTULA,并给出其性能的理论保障。具体而言,在非渐近条件下建立了KL散度的收敛界,收敛速率可达2−ε(ε>0),为现有文献中最优。由此可推导出在Wasserstein-2距离下的改进误差界,进一步用于解决相关优化问题的非渐近保证。通过在高维双阱势分布采样及神经网络优化问题上的应用,验证了该算法的有效性与理论可证性。

原文摘要 · Abstract (English)

Motivated by applications in deep learning, where the global Lipschitz continuity condition is often not satisfied, we examine the problem of sampling from distributions with super-linearly growing log-gradients. We propose a novel tamed Langevin dynamics-based algorithm, called kTULA, to solve the aforementioned sampling problem, and provide a theoretical guarantee for its performance. More precisely, we establish a non-asymptotic convergence bound in Kullback-Leibler (KL) divergence with the best-known rate of convergence equal to $2-\overlineε$, $\overlineε>0$, which significantly improves relevant results in existing literature. This enables us to obtain an improved non-asymptotic error bound in Wasserstein-2 distance, which can be used to further derive a non-asymptotic guarantee for kTULA to solve the associated optimization problems. To illustrate the applicability of kTULA, we apply the proposed algorithm to the problem of sampling from a high-dimensional double-well potential distribution and to an optimization problem involving a neural network. We show that our main results can be used to provide theoretical guarantees for the performance of kTULA.

采样算法Langevin优化概率采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。