用声学特征预测语音质量评分,无需人工打分。
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
- 提取音频的通用隐空间特征,实现端到端预测。
- 在MSE、LCC等指标上达到当前最优效果。
- 模型小、通用性强,适合快速筛选语音合成模型。
语音质量评估在文本转语音(TTS)或语音转换模型选择中至关重要。现有客观指标如PESQ、POLQA、STOI无法有效筛选最佳模型;而主观指标均值意见分(MOS)虽可靠但耗时费力。为此,我们提出SALF-MOS模型,一种小型、端到端、高度泛化且可扩展的语音质量评分预测模型,可输出5分制分数。通过堆叠卷积层提取音频样本的隐空间特征,在均方误差(MSE)、线性一致性相关系数(LCC)、斯皮尔曼等级相关系数(SRCC)和肯德尔等级相关系数(KTAU)上取得当前最优表现。
原文摘要 · Abstract (English)
Speech quality assessment is a critical process in selecting text-to-speech synthesis (TTS) or voice conversion models. Evaluation of voice synthesis can be done using objective metrics or subjective metrics. Although there are many objective metrics like the Perceptual Evaluation of Speech Quality (PESQ), Perceptual Objective Listening Quality Assessment (POLQA) or Short-Time Objective Intelligibility (STOI) but none of them is feasible in selecting the best model. On the other hand subjective metric like Mean Opinion Score is highly reliable but it requires a lot of manual efforts and are time-consuming. To counter the issues in MOS Evaluation, we have developed a novel model, Speaker Agnostic Latent Features (SALF)-Mean Opinion Score (MOS) which is a small-sized, end-to-end, highly generalized and scalable model for predicting MOS score on a scale of 5. We use the sequences of convolutions and stack them to get the latent features of the audio samples to get the best state-of-the-art results based on mean squared error (MSE), Linear Concordance Correlation coefficient (LCC), Spearman Rank Correlation Coefficient (SRCC) and Kendall Rank Correlation Coefficient (KTAU).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。