针对台湾国语多音字难题,打造更逼真的语音合成系统
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
- 用LLM+流匹配模型增强发音控制,精准处理多音字歧义
- 在通用和中英混用场景下均表现优于现有系统
- 适合需要高保真语音合成的本地化应用开发者
我们提出BreezyVoice,一种专为台湾国语定制的文本转语音系统,重点解决该语言中多音字辨识难题。基于CosyVoice架构,引入S³分词器、大语言模型(LLM)、最优传输条件流匹配模型(OT-CFM)及音素预测模型,生成接近真人语调的自然语音。评估显示,BreezyVoice在通用与中英混写场景下均表现优异,具备强鲁棒性与高保真度。同时,我们探讨了长尾说话人建模与多音字歧义消解的泛化挑战,方法显著提升性能,并为神经编解码语音合成系统的设计提供关键洞见。
原文摘要 · Abstract (English)
We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon CosyVoice, we incorporate a $S^{3}$ tokenizer, a large language model (LLM), an optimal-transport conditional flow matching model (OT-CFM), and a grapheme to phoneme prediction model, to generate realistic speech that closely mimics human utterances. Our evaluation demonstrates BreezyVoice's superior performance in both general and code-switching contexts, highlighting its robustness and effectiveness in generating high-fidelity speech. Additionally, we address the challenges of generalizability in modeling long-tail speakers and polyphone disambiguation. Our approach significantly enhances performance and offers valuable insights into the workings of neural codec TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。