用英语语音模型微调出能流利切换印度语的超拟真合成系统
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
- 用英语语音模型在印度语数据上微调,实现跨语言自然发音
- 11种印度语言中表现接近真人水平,支持跨语言流畅说话
- 适合低资源语言语音合成研究者,尤其关注零样本生成
当英语童话语音模型在印度语言上进行微调时会发生什么?我们评估了英语F5-TTS模型在11种印度语言上的适应能力,涵盖多语言流畅度、语音克隆、风格克隆和混语生成。对比三种方式:(i) 从头训练,(ii) 仅用印度语数据微调,(iii) 同时使用印度语与英语数据以避免遗忘。结果显示,仅用印度语数据微调效果最佳,生成的IN-F5模型接近人类水平的多语言表达,使一种语言使用者(如奥里亚语)可自然说出另一种语言(如印地语)。实验表明,英语预训练有助于低资源语音合成达到人类水准。我们还研究了数据受限场景,提出计算最优策略。最后,通过人机协同生成合成数据,成功实现对博杰布尔语和图鲁语等未见语言的零资源语音合成。
原文摘要 · Abstract (English)
What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing. We compare: (i) training from scratch, (ii) fine-tuning English F5 on Indian data, and (iii) fine-tuning on both Indian and English data to prevent forgetting. Fine-tuning with only Indian data proves most effective and the resultant IN-F5 is a near-human polyglot; that enables speakers of one language (e.g., Odia) to fluently speak in another (e.g., Hindi). Our results show English pretraining aids low-resource TTS in reaching human parity. To aid progress in other low-resource languages, we study data-constrained setups and arrive at a compute optimal strategy. Finally, we show IN-F5 can synthesize unseen languages like Bhojpuri and Tulu using a human-in-the-loop approach for zero-resource TTS via synthetic data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。