arXiv:2503.07021cs.LGstat.ML2025-03被引 1

用可学习参数简化能量模型训练,无需复杂采样

Learning Energy-Based Models by Self-normalising the Likelihood

  • 引入可学习的归一化常数,构建自归一化似然目标
  • 在指数族分布下目标函数凹,支持直接随机梯度优化
  • 仅需粗略采样即可训练,性能优于传统方法

基于能量的模型(EBM)的最大似然训练因归一化常数不可计算而困难。传统方法依赖昂贵的马尔可夫链蒙特卡洛(MCMC)采样估算归一化常数的梯度。本文提出一种新目标——自归一化对数似然(SNL),相比常规对数似然仅增加一个可学习参数来表示归一化常数。SNL是对数似然的下界,其最优解同时对应模型参数和归一化常数的最大似然估计。我们证明,在指数族分布下,SNL目标函数关于模型参数是凹的。与常规对数似然不同,SNL可通过从粗糙提议分布中采样直接使用随机梯度优化。我们在多种密度估计任务及回归用的EBM上验证了该方法的有效性。结果表明,该方法实现更简单、调参更少,且性能超越现有技术。

原文摘要 · Abstract (English)

Training an energy-based model (EBM) with maximum likelihood is challenging due to the intractable normalisation constant. Traditional methods rely on expensive Markov chain Monte Carlo (MCMC) sampling to estimate the gradient of logartihm of the normalisation constant. We propose a novel objective called self-normalised log-likelihood (SNL) that introduces a single additional learnable parameter representing the normalisation constant compared to the regular log-likelihood. SNL is a lower bound of the log-likelihood, and its optimum corresponds to both the maximum likelihood estimate of the model parameters and the normalisation constant. We show that the SNL objective is concave in the model parameters for exponential family distributions. Unlike the regular log-likelihood, the SNL can be directly optimised using stochastic gradient techniques by sampling from a crude proposal distribution. We validate the effectiveness of our proposed method on various density estimation tasks as well as EBMs for regression. Our results show that the proposed method, while simpler to implement and tune, outperforms existing techniques.

能量模型归一化生成模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。