无需目标说话人多语言数据,实现自然的跨语言语音转换。
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
- 通过跨语言微调,让单语模型也能保留源语音口音。
- 在英、西、法、中四语种测试中,语音可懂度和质量均优于现有方法。
- 适合语音克隆、多语言语音合成等需要保留口音的应用场景。
尽管近年来取得了显著进展,但构建人工多语种语音仍是一项挑战。本文研究了自监督学习在语音转换中的应用,以生成自然的多语种语音。提出一种新型的跨语言任意到单一语音转换系统,能够在不依赖目标说话人多语言数据的情况下保留源语音口音。此外,我们提出一种新颖的跨语言微调策略,进一步提升口音还原度并减少训练数据需求。对英语、西班牙语、法语和汉语普通话的客观与主观评估表明,该方法优于现有最先进方法,显著提升了跨语言场景下的语音可懂度与整体质量。音频样例见 https://giuseppe-ruggiero.github.io/a2o-vc-demo/
原文摘要 · Abstract (English)
The creation of artificial polyglot voices remains a challenging task, despite considerable progress in recent years. This paper investigates self-supervised learning for voice conversion to create native-sounding polyglot voices. We introduce a novel cross-lingual any-to-one voice conversion system that is able to preserve the source accent without the need for multilingual data from the target speaker. In addition, we show a novel cross-lingual fine-tuning strategy that further improves the accent and reduces the training data requirements. Objective and subjective evaluations with English, Spanish, French and Mandarin Chinese confirm that our approach improves on state-of-the-art methods, enhancing the speech intelligibility and overall quality of the converted speech, especially in cross-lingual scenarios. Audio samples are available at https://giuseppe-ruggiero.github.io/a2o-vc-demo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。