让语音合成跨越方言,自然模仿不同口音的说话方式。
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
- 用多方言音素级BERT预测目标方言的语调特征
- 合成语音在跨方言任务中自然度显著提升
- 适合需要多口音语音助手的研究与开发
我们探索跨方言文本到语音(CD-TTS)任务,旨在合成非母语方言的语音,尤其针对语调重音语言。该任务对开发能自然与各地人群交流的语音助手至关重要。本文提出一种新型TTS模型,包含三个子模块:首先训练一个基础TTS模型,通过音素级语调潜在变量(ALVs)从文本生成方言语音;其次训练一个ALV预测器,利用新提出的多方言音素级BERT,从输入文本预测目标方言的语调特征;最后通过多方言语音合成实验,对比基于传统方言TTS方法的基线模型。结果表明,本模型在跨方言语音合成任务中显著提升了合成语音的方言自然度。
原文摘要 · Abstract (English)
We explore cross-dialect text-to-speech (CD-TTS), a task to synthesize learned speakers' voices in non-native dialects, especially in pitch-accent languages. CD-TTS is important for developing voice agents that naturally communicate with people across regions. We present a novel TTS model comprising three sub-modules to perform competitively at this task. We first train a backbone TTS model to synthesize dialect speech from a text conditioned on phoneme-level accent latent variables (ALVs) extracted from speech by a reference encoder. Then, we train an ALV predictor to predict ALVs tailored to a target dialect from input text leveraging our novel multi-dialect phoneme-level BERT. We conduct multi-dialect TTS experiments and evaluate the effectiveness of our model by comparing it with a baseline derived from conventional dialect TTS methods. The results show that our model improves the dialectal naturalness of synthetic speech in CD-TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。