arXiv:2505.13771cs.SDcs.LG2025-05

提出新方法提升能量模型的梯度学习,让语音合成更稳定高效。

Score-Based Training for Energy-Based TTS Models

  • 基于分数的训练准则,直接优化一阶优化所需的梯度方向。
  • 在语音合成任务中,相比NCE和SSM,生成质量显著提升。
  • 适合研究能量模型与扩散模型融合的语音生成方向者阅读。

噪声对比估计(NCE)是训练具有不可计算归一化项的能量模型(EBM)的常用方法,其核心思想是通过比较参考样本与噪声样本的非归一化对数似然来学习,从而避免显式计算归一化项。然而,NCE严重依赖噪声样本的质量。近年来,切片分数匹配(SSM)因与扩散模型(DM)密切相关而受到关注,其通过学习随机方向上对数似然投影的分布来估计得分。但无论是NCE还是SSM,均忽略了对数似然函数的具体形式,这对使用一阶优化进行推理的EBM和DM而言是潜在问题。本文提出一种新准则,旨在学习更适合一阶优化方案的得分。实验对比了多种方法在训练EBM上的表现。

原文摘要 · Abstract (English)

Noise contrastive estimation (NCE) is a popular method for training energy-based models (EBM) with intractable normalisation terms. The key idea of NCE is to learn by comparing unnormalised log-likelihoods of the reference and noisy samples, thus avoiding explicitly computing normalisation terms. However, NCE critically relies on the quality of noisy samples. Recently, sliced score matching (SSM) has been popularised by closely related diffusion models (DM). Unlike NCE, SSM learns a gradient of log-likelihood, or score, by learning distribution of its projections on randomly chosen directions. However, both NCE and SSM disregard the form of log-likelihood function, which is problematic given that EBMs and DMs make use of first-order optimisation during inference. This paper proposes a new criterion that learns scores more suitable for first-order schemes. Experiments contrasts these approaches for training EBMs.

语音合成能量模型分数匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。