用专家混合机制提升TTS对陌生描述的理解能力
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
- 通过模态专用专家网络增强预训练语言模型的语音理解能力
- 在陌生描述测试集上优于最先进商业产品
- 适合需要强泛化能力的语音合成场景
基于描述的文本转语音(TTS)模型在训练域内描述上表现良好,但在真实应用中,用户生成的多样化描述常引入大量域外输入,挑战系统文本理解能力。为此,我们提出MoE-TTS,一种旨在增强域外描述理解的描述式TTS模型。MoE-TTS采用基于模态的专家混合(MoE)方法,在保持预训练文本大语言模型(LLM)冻结的前提下,为其注入一组适配语音模态的专用权重。该方法使模型能有效利用预训练知识与文本理解能力。实验表明:首先,即使最先进的闭源商用产品也难以应对精心设计的域外描述测试集;其次,MoE-TTS在生成更准确反映描述内容的语音方面表现更优。建议听众访问演示页面 https://welkinyang.github.io/MoE-TTS/。
原文摘要 · Abstract (English)
Description-based text-to-speech (TTS) models exhibit strong performance on in-domain text descriptions, i.e., those encountered during training. However, in real-world applications, the diverse range of user-generated descriptions inevitably introduces numerous out-of-domain inputs that challenge the text understanding capabilities of these systems. To address this issue, we propose MoE-TTS, a description-based TTS model designed to enhance the understanding of out-of-domain text descriptions. MoE-TTS employs a modality-based mixture-of-experts (MoE) approach to augment a pre-trained textual large language model (LLM) with a set of specialized weights adapted to the speech modality while maintaining the original LLM frozen during training. This approach allows MoE-TTS to effectively leverage the pre-trained knowledge and text understanding abilities of textual LLMs. Our experimental results indicate that: first, even the most advanced closed-source commercial products can be challenged by carefully designed out-of-domain description test sets; second, MoE-TTS achieves superior performance in generating speech that more accurately reflects the descriptions. We encourage readers to listen to the demos at https://welkinyang.github.io/MoE-TTS/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。