用图像模型分析语音谱图,提升高质量合成语音自然度评分预测精度。
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
- 将预训练图像模型用于语音谱图特征提取,捕捉合成语音差异。
- 在16项评测中7项第一、9项第二,显著领先第三名。
- 适合语音质量评估、语音合成优化方向的研究者参考。
我们提出了T05系统,用于2024年语音自然度评分挑战赛(VoiceMOS Challenge, VMC 2024)的第1赛道,该赛道专注于高质量合成语音自然度均值意见分(MOS)的准确预测。除使用预训练自监督学习(SSL)语音特征提取器外,系统还引入了预训练图像特征提取器,以捕捉语音谱图中合成语音的差异。我们首先分别训练基于SSL特征或谱图特征的MOS预测器,随后通过融合两种特征对两个预测器进行微调,提升预测性能。在VMC 2024第1赛道中,T05系统在16项评估指标中取得7项第一、9项第二,显著优于排名第三及以下的系统。我们还报告了消融实验结果,以探究系统关键因素。
原文摘要 · Abstract (English)
We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。