arXiv:2409.03717cs.SDcs.AI2024-09被引 5

用不到1000小时数据实现高质量语音合成,效果超越主流自回归模型。

Sample-Efficient Diffusion for Text-To-Speech Synthesis

  • 在预训练音频自编码器的潜在空间中,用新型U-Audio Transformer进行扩散建模。
  • 仅用不到1000小时数据训练,生成语音清晰度超过当前自回归模型VALL-E。
  • 适合低资源语音合成场景,尤其适用于数据稀缺的垂直领域应用。

本文提出样本高效的语音扩散模型SESD,通过潜在空间中的扩散机制,在小规模数据下实现高效语音合成。其核心是新型扩散架构U-Audio Transformer(U-AT),能高效处理长序列,并在预训练音频自编码器的潜在空间中运行。在字符感知语言模型表示条件下,SESD仅需少于1000小时语音数据即可训练,生成语音的可懂度反而高于当前最先进的自回归模型VALL-E,且训练数据不足其2%。该方法在数据效率和语音质量之间实现了显著平衡。

原文摘要 · Abstract (English)

This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Audio Transformer (U-AT), that efficiently scales to long sequences and operates in the latent space of a pre-trained audio autoencoder. Conditioned on character-aware language model representations, SESD achieves impressive results despite training on less than 1k hours of speech - far less than current state-of-the-art systems. In fact, it synthesizes more intelligible speech than the state-of-the-art auto-regressive model, VALL-E, while using less than 2% the training data.

语音合成扩散模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。