用少量数据实现零样本多任务语音合成,还能保持语气自然。
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
- 基于声学理论解耦声源与滤波器特征,用提示词建模
- 仅需极少训练数据,零样本语音合成效果接近大数据模型
- 创新的韵律建模提升语音自然度,适合小样本场景
文本到语音(TTS)系统通过增加训练数据量在零样本语音合成方面取得显著进展。然而,这类系统存在训练数据需求量大、成本高,且常忽略韵律相似性等问题。为此,我们提出MultiVerse,一种能够在零样本和跨语言条件下进行语音合成或风格迁移的零样本多任务TTS系统。该系统所需训练数据远少于传统数据驱动方法。为在数据有限时仍保持零样本性能,我们采用基于源-滤波器理论的解耦机制,利用提示词建模滤波器相关与声源相关表征。为进一步提升韵律相似性,我们结合提示词驱动的自回归与非自回归方法进行韵律建模。实验表明,MultiVerse在零样本多任务TTS上表现优异,不仅在极少量数据下达到与大数据模型相当的合成效果,还显著优于其他同规模数据训练的零样本TTS系统。特别地,其创新的韵律建模技术显著提升了生成语音与提示语的韵律相似性。音频样例见:https://nc-ai.github.io/speech/publications/multiverse/index.html
原文摘要 · Abstract (English)
Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of training data, which increases costs, and often overlook prosody similarity. To address these issues, we propose MultiVerse, a zero-shot multi-task TTS system that is able to perform TTS or speech style transfer in zero-shot and cross-lingual conditions. MultiVerse requires much less training data than traditional data-driven approaches. To ensure zero-shot performance even with limited data, we leverage source-filter theory-based disentanglement, utilizing the prompt for modeling filter-related and source-related representations. Additionally, to further enhance prosody similarity, we adopt a prosody modeling approach combining prompt-based autoregressive and non-autoregressive methods. Evaluations demonstrate the remarkable zero-shot multi-task TTS performance of MultiVerse and show that MultiVerse not only achieves zero-shot TTS performance comparable to data-driven TTS systems with much less data, but also significantly outperforms other zero-shot TTS systems trained with the same small amount of data. In particular, our novel prosody modeling technique significantly contributes to MultiVerse's ability to generate speech with high prosody similarity to the given prompts. Our samples are available at https://nc-ai.github.io/speech/publications/multiverse/index.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。