arXiv:2409.09351eess.AScs.SD2024-09被引 13

E1 TTS 一步生成语音,速度快且音质自然。

E1 TTS: Simple and Fast Non-Autoregressive TTS

  • 基于去噪扩散预训练与分布匹配蒸馏,无需对齐文本音频
  • 单次前向计算即可生成完整语音,推理效率高
  • 零样本迁移能力强,适合快速部署语音合成系统

本文提出高效非自回归零样本文语转换系统 E1 TTS,基于去噪扩散预训练与分布匹配蒸馏。训练过程简单,无需显式对齐文本与音频序列。推理仅需对每个语音片段进行一次神经网络前向计算,效率极高。尽管采样高效,E1 TTS 在语音自然度与说话人相似性方面达到与多个强基线模型相当的水平。音频样例可访问 http://e1tts.github.io/。

原文摘要 · Abstract (English)

This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is straightforward; it does not require explicit monotonic alignment between the text and audio pairs. The inference of E1 TTS is efficient, requiring only one neural network evaluation for each utterance. Despite its sampling efficiency, E1 TTS achieves naturalness and speaker similarity comparable to various strong baseline models. Audio samples are available at http://e1tts.github.io/ .

语音合成非自回归扩散模型快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。