arXiv:2502.01084cs.LGcs.SD2025-02ICLR被引 6

用连续表示替代量化,用概率模型实现高效语音合成。

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

  • 用VAE+GMM构建连续自回归语音模型,避免传统量化复杂度。
  • 参数量仅为VALL-E的10.3%,主观与客观评测均更优。
  • 适合追求轻量高效语音合成的研究者与开发者。

我们提出一种新型自回归语音合成建模方法,结合变分自编码器(VAE)与多模态潜在空间,采用高斯混合模型(GMM)作为条件概率分布。不同于依赖残差向量量化的现有方法,本模型利用VAE潜在空间中的连续语音表示,显著简化训练与推理流程。同时引入随机单调对齐机制,强制严格单调对齐。实验表明,该方法在主观与客观评估中均显著优于当前最先进的自回归模型VALL-E,仅需其10.3%的参数量即可达成更优效果。这证明了连续语音语言模型在效率上具有超越现有量化基模型的潜力。样本音频见 https://tinyurl.com/gmm-lm-tts。

原文摘要 · Abstract (English)

We propose a novel autoregressive modeling approach for speech synthesis, combining a variational autoencoder (VAE) with a multi-modal latent space and an autoregressive model that uses Gaussian Mixture Models (GMM) as the conditional probability distribution. Unlike previous methods that rely on residual vector quantization, our model leverages continuous speech representations from the VAE's latent space, greatly simplifying the training and inference pipelines. We also introduce a stochastic monotonic alignment mechanism to enforce strict monotonic alignments. Our approach significantly outperforms the state-of-the-art autoregressive model VALL-E in both subjective and objective evaluations, achieving these results with only 10.3\% of VALL-E's parameters. This demonstrates the potential of continuous speech language models as a more efficient alternative to existing quantization-based speech language models. Sample audio can be found at https://tinyurl.com/gmm-lm-tts.

语音合成连续表示自回归轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。