arXiv:2412.08117cs.SDcs.AI2024-12被引 3

用潜在扩散模型生成语音,速度更快、质量更高。

LatentSpeech: Latent Diffusion for Text-To-Speech Generation

  • 在潜在空间中用扩散模型生成语音,压缩数据维度至5%
  • 比现有模型降低25%词错误率,降低24%梅尔倒谱失真
  • 适合追求高效高质语音合成的研究与开发者

基于扩散的生成式AI在计算机视觉和自然语言处理中表现优异,但在语音生成领域仍不成熟。主流文本转语音系统依赖频谱空间中的梅尔频谱图(Mel-Spectrograms),因频谱稀疏导致计算开销大。本文提出LatentSpeech,一种基于潜在扩散模型的新颖文本转语音方法。通过使用潜在嵌入作为中间表示,将目标维度降至梅尔频谱所需尺寸的5%,显著简化编码器与声码器的处理流程,实现高效高质量语音生成。这是首个将潜在扩散模型应用于文本转语音的工作,有效提升了生成语音的准确性和自然度。在基准数据集上的实验表明,相比现有模型,LatentSpeech在词错误率上提升25%,梅尔倒谱失真降低24%;额外训练数据下,两项指标进一步提升至49.5%和26%。结果证明该方法具有推动语音合成技术进步的巨大潜力。

原文摘要 · Abstract (English)

Diffusion-based Generative AI gains significant attention for its superior performance over other generative techniques like Generative Adversarial Networks and Variational Autoencoders. While it has achieved notable advancements in fields such as computer vision and natural language processing, their application in speech generation remains under-explored. Mainstream Text-to-Speech systems primarily map outputs to Mel-Spectrograms in the spectral space, leading to high computational loads due to the sparsity of MelSpecs. To address these limitations, we propose LatentSpeech, a novel TTS generation approach utilizing latent diffusion models. By using latent embeddings as the intermediate representation, LatentSpeech reduces the target dimension to 5% of what is required for MelSpecs, simplifying the processing for the TTS encoder and vocoder and enabling efficient high-quality speech generation. This study marks the first integration of latent diffusion models in TTS, enhancing the accuracy and naturalness of generated speech. Experimental results on benchmark datasets demonstrate that LatentSpeech achieves a 25% improvement in Word Error Rate and a 24% improvement in Mel Cepstral Distortion compared to existing models, with further improvements rising to 49.5% and 26%, respectively, with additional training data. These findings highlight the potential of LatentSpeech to advance the state-of-the-art in TTS technology

语音合成扩散模型潜在空间高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。