arXiv:2608.06283cs.LGmath.OC2026-08

针对非凸、梯度超线性增长的复杂分布,提出稳定高效的采样算法。

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

论文配图:The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
图 1 · 摘自论文原文
  • 直接使用次梯度,结合压制技术实现稳定离散化。
  • 在Wasserstein-2距离下给出显式收敛界,优于现有方法。
  • 适用于大模型预训练,性能媲美优化器且有理论保障。

我们研究目标分布势函数同时具有非光滑性、超线性梯度增长和非凸性的采样问题。提出次梯度压制无调整Langevin算法(SG-TULA),直接基于次梯度操作,无需计算代价高的平滑处理。为应对超线性情形,采用压制技术构建稳定显式方案。在Wasserstein-2距离下导出非渐近收敛界,所有常数显式依赖于维度与逆温度,优于当前已知的次梯度类Langevin算法。进一步提供关联优化问题的过拟合风险估计。验证了GPT-2系列大模型正则化预训练势函数的假设条件,常数明确;并证明改进的坐标压缩版SG-TULA在预训练中可媲美微调后的AdamW与Muon,且前者无类似非渐近保证。

原文摘要 · Abstract (English)

We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.

采样算法非凸优化大模型预训练理论保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。