arXiv:2505.19087cs.LGstat.ML2025-05NeurIPS被引 2

温度决定模型泛化能力,无需依赖训练时长或混合性。

Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes

  • 用温度控制的Langevin动态,泛化误差可被温度与初始损失统一约束。
  • 在任意训练时刻,泛化误差上界为√(β·E[L(θ₀)] + log(1/δ))/√N,与维度、梯度无关。
  • 适用于过参数化模型,尤其适合关注泛化机制的研究者。

我们分析了使用马尔可夫随机训练算法(如带正温度β⁻¹的Langevin动力学)训练过参数化模型时的泛化差距。该算法在极小步长下进行梯度下降,并加入方差为β⁻¹的高斯噪声,且轻微正则化或有界。我们证明,在任意训练时刻,泛化误差以概率1−δ满足上界√(β𝔼L(θ₀) + log(1/δ))/√N,其中N为样本量,𝔼L(θ₀)=O(1)。该结果不依赖于训练时间、混合性、维度、梯度范数或损失函数特性。此结论源于对具有吉布斯型平稳分布的任意马尔可夫过程的通用分析,其证明基于广义热力学第二定律所隐含的初始分布偏移有界性。

原文摘要 · Abstract (English)

We analyze the generalization gap (gap between the training and test errors) when training a potentially over-parametrized model using a Markovian stochastic training algorithm, initialized from some distribution $θ_0 \sim p_0$. We focus on Langevin dynamics with a positive temperature $β^{-1}$, i.e. gradient descent on a training loss $L$ with infinitesimal step size, perturbed with $β^{-1}$-variances Gaussian noise, and lightly regularized or bounded. There, we bound the generalization gap, at any time during training, by $\sqrt{(β\mathbb{E} L (θ_0) + \log(1/δ))/N}$ with probability $1-δ$ over the dataset, where $N$ is the sample size, and $\mathbb{E} L (θ_0) =O(1)$ with standard initialization scaling. In contrast to previous guarantees, we have no dependence on either training time or reliance on mixing, nor a dependence on dimensionality, gradient norms, or any other properties of the loss or model. This guarantee follows from a general analysis of any Markov process-based training that has a Gibbs-style stationary distribution. The proof is surprisingly simple, once we observe that the marginal distribution divergence from initialization remains bounded, as implied by a generalized second law of thermodynamics.

泛化分析Langevin动力学马尔可夫过程温度调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。