arXiv:2410.16048eess.AS2024-10被引 8

用分词扩散模型直接生成连续语音表示,零样本合成更清晰。

Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion

  • 每帧用扩散过程迭代优化连续语音特征
  • 语音可懂度更高,音质与说话人相似度媲美真音频
  • 适合追求高可懂度的零样本语音合成场景

我们提出SALAD,一种在连续语音表示上运行的零样本端到端语音合成自回归模型。SALAD采用逐词扩散过程,对下一时间步的连续表示进行精炼与预测。我们将其与离散版本的SALAD以及公开可用的零样本语音合成系统进行对比,并全面分析了离散与连续建模技术的差异。结果表明,SALAD在保持真实音频的语音质量与说话人相似度的同时,显著提升了语音可懂度。

原文摘要 · Abstract (English)

We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio.

语音合成扩散模型连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。