用不到1000小时数据实现高质量语音合成,效果超越主流自回归模型。
Sample-Efficient Diffusion for Text-To-Speech Synthesis
- 在预训练音频自编码器的潜在空间中,用新型U-Audio Transformer进行扩散建模。
- 仅用不到1000小时数据训练,生成语音清晰度超过当前自回归模型VALL-E。
- 适合低资源语音合成场景,尤其适用于数据稀缺的垂直领域应用。
本文提出样本高效的语音扩散模型SESD,通过潜在空间中的扩散机制,在小规模数据下实现高效语音合成。其核心是新型扩散架构U-Audio Transformer(U-AT),能高效处理长序列,并在预训练音频自编码器的潜在空间中运行。在字符感知语言模型表示条件下,SESD仅需少于1000小时语音数据即可训练,生成语音的可懂度反而高于当前最先进的自回归模型VALL-E,且训练数据不足其2%。该方法在数据效率和语音质量之间实现了显著平衡。
原文摘要 · Abstract (English)
This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Audio Transformer (U-AT), that efficiently scales to long sequences and operates in the latent space of a pre-trained audio autoencoder. Conditioned on character-aware language model representations, SESD achieves impressive results despite training on less than 1k hours of speech - far less than current state-of-the-art systems. In fact, it synthesizes more intelligible speech than the state-of-the-art auto-regressive model, VALL-E, while using less than 2% the training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。