arXiv:2409.09311eess.AScs.SD2024-09中稿 · ICASSP 2025被引 2

通过稳定共振峰生成提升扩散模型零样本语音合成的发音鲁棒性

Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation

  • 引入声源-滤波器理论,设计稳定共振峰生成架构
  • 在未见说话人上实现更高发音准确率与自然度
  • 适合关注语音合成发音质量与可扩展性的研究者

扩散模型在文本转语音(TTS)任务中取得了显著进展,尤其在零样本场景下表现优异。尽管现有工作聚焦于推理速度与音质的权衡,我们发现扩散过程导致的发音不稳定问题被忽视。通过初步研究,我们观察到扩散过程引发的发音不稳定性。为此,提出StableForm-TTS,一种新型零样本语音合成框架,在保持扩散建模优势的同时提升发音鲁棒性。首次将声源-滤波器理论应用于扩散TTS,设计了精细的共振峰生成架构。在未见说话人上的实验表明,该模型在发音准确率和自然度上优于当前最先进方法,且说话人相似性相当。此外,模型在数据量与模型规模增大时仍具良好可扩展性。音频样例已公开:https://deepbrainai-research.github.io/stableformtts/

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of diffusion models. However, we find a critical mispronunciation issue is being overlooked. Our preliminary study reveals the unstable pronunciation resulting from the diffusion process. Based on this observation, we introduce StableForm-TTS, a novel zero-shot speech synthesis framework designed to produce robust pronunciation while maintaining the advantages of diffusion modeling. By pioneering the adoption of source-filter theory in diffusion TTS, we propose an elaborate architecture for stable formant generation. Experimental results on unseen speakers show that our model outperforms the state-of-the-art method in terms of pronunciation accuracy and naturalness, with comparable speaker similarity. Moreover, our model demonstrates effective scalability as both data and model sizes increase. Audio samples are available online: https://deepbrainai-research.github.io/stableformtts/.

语音合成扩散模型零样本发音鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。