用5000小时数据训练语音克隆模型,让标准德语变瑞士方言语音。
Voice Adaptation for Swiss German
- 用播客数据自动标注方言,构建5000小时弱标签训练集
- 模型在人评和机评中表现良好,CMOS达-0.28,SMOS达3.8
- 为低资源方言语音合成提供可复现的适配方案
本文研究语音适配模型在瑞士德语方言中的表现,即把标准德语文本转换为瑞士方言语音。我们预处理了一个大规模瑞士播客数据集,通过自动转录和方言分类标注,获得约5000小时弱标签训练数据。在该数据集上微调XTTSv2模型,结果显示其在人工与自动评估中均表现优异,能准确生成目标方言语音。本工作推动了语音克隆技术向非主流语言的适配。最终模型达到最高CMOS分数-0.28和SMOS分数3.8。
原文摘要 · Abstract (English)
This work investigates the performance of Voice Adaptation models for Swiss German dialects, i.e., translating Standard German text to Swiss German dialect speech. For this, we preprocess a large dataset of Swiss podcasts, which we automatically transcribe and annotate with dialect classes, yielding approximately 5000 hours of weakly labeled training material. We fine-tune the XTTSv2 model on this dataset and show that it achieves good scores in human and automated evaluations and can correctly render the desired dialect. Our work shows a step towards adapting Voice Cloning technology to underrepresented languages. The resulting model achieves CMOS scores of up to -0.28 and SMOS scores of 3.8.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。