arXiv:2409.01217cs.CLcs.LG2024-09被引 5

用多语言预训练提升低资源语音合成的自然度与可懂度

A multilingual training strategy for low resource Text to Speech

  • 基于社交媒体数据,通过有选择地融合多语言语料进行预训练
  • 相比单语训练,多语言预训练使语音可懂度和自然度显著提升
  • 适合低资源语言语音合成研究者参考,尤其关注跨语言迁移

近期神经文本到语音(TTS)技术已能生成高质量语音,但其依赖大量标注数据,难以低成本扩展至所有语言,尤其是低资源语言。本文探讨两个问题:一是能否利用社交媒体数据构建小规模TTS数据集,二是跨语言迁移学习在低资源语言上是否有效。重点评估多语言建模作为单语数据替代方案的可行性。通过研究外语文本数据的选择与融合策略,发现经过合理筛选的多语言预训练显著优于单语预训练,在提升目标语言语音的可懂度与自然度方面表现更优。实验表明,使用多语言预训练配合源语言选择策略,可在低资源场景下实现更优的语音合成效果。

原文摘要 · Abstract (English)

Recent speech technologies have led to produce high quality synthesised speech due to recent advances in neural Text to Speech (TTS). However, such TTS models depend on extensive amounts of data that can be costly to produce and is hardly scalable to all existing languages, especially that seldom attention is given to low resource languages. With techniques such as knowledge transfer, the burden of creating datasets can be alleviated. In this paper, we therefore investigate two aspects; firstly, whether data from social media can be used for a small TTS dataset construction, and secondly whether cross lingual transfer learning (TL) for a low resource language can work with this type of data. In this aspect, we specifically assess to what extent multilingual modeling can be leveraged as an alternative to training on monolingual corporas. To do so, we explore how data from foreign languages may be selected and pooled to train a TTS model for a target low resource language. Our findings show that multilingual pre-training is better than monolingual pre-training at increasing the intelligibility and naturalness of the generated speech.

语音合成多语言低资源迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。