arXiv:2510.13438stat.MLcs.LG2025-10NeurIPS被引 2

对比散度算法在特定条件下可逼近最优学习速率。

Near-Optimality of Contrastive Divergence Algorithms

  • 在正则条件下,对比散度实现参数化收敛速度。
  • 在不同数据分批策略下均达到近似最优性能。
  • 适合研究生成模型训练效率的学者参考。

我们对未归一化模型的对比散度(CD)算法进行了非渐近分析。尽管已有研究证明,对于指数族分布,CD 迭代的渐近收敛速率为 $O(n^{-1/3})$,但本文在一定正则性假设下表明,CD 可达到 $O(n^{-1/2})$ 的参数化收敛速率。分析覆盖了全在线和小批量等多种数据分批方案,并进一步证明,CD 的渐近方差接近 Cramér-Rao 下界,具有近似最优性。

原文摘要 · Abstract (English)

We perform a non-asymptotic analysis of the contrastive divergence (CD) algorithm, a training method for unnormalized models. While prior work has established that (for exponential family distributions) the CD iterates asymptotically converge at an $O(n^{-1 / 3})$ rate to the true parameter of the data distribution, we show, under some regularity assumptions, that CD can achieve the parametric rate $O(n^{-1 / 2})$. Our analysis provides results for various data batching schemes, including the fully online and minibatch ones. We additionally show that CD can be near-optimal, in the sense that its asymptotic variance is close to the Cramér-Rao lower bound.

对比散度生成模型收敛速率统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。