用适配器实现轻量级语音合成中未见过的说话人和语言迁移。
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
- 通过适配器让预训练模型学习新说话人或新语言特征。
- 在无目标语音样本情况下,仍能生成接近母语口音的语音。
- 适合需要多语言、多说话人快速部署的轻量级语音系统开发者。
本文研究在轻量级文本转语音(TTS)系统中,通过适配器实现跨语言语音合成。重点比较了未见过的说话人与语言适应任务,目标是在目标语言中合成目标说话人的声音,而该说话人在此语言中无录音数据。客观评估表明,适配器能有效学习语言特异性和说话人特异性信息,使预训练模型学会从未见过的说话人身份或语言,同时避免原模型说话人或语言信息的灾难性遗忘。此外,为衡量生成语音的口音自然度,提出并验证了一种受第二语言(L2)发音错误检测技术启发的客观指标。论文还分析了适配器位置、配置及训练说话人数对性能的影响。
原文摘要 · Abstract (English)
In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。