用连续评分建模提升语音质量评估准确率
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
- 将人工评分视为连续值的离散化结果,构建量化分布模型
- 新方法使MOSNet预测误差降低12.3%,在LIBRISSR数据集上验证
- 适合做语音合成与语音转换的质量评估研究者参考
语音质量评估(SQA)旨在不依赖耗时的听觉问卷的情况下评估语音样本质量。近年来的研究聚焦于训练基于神经网络的SQA模型,以预测文本转语音或语音转换系统生成语音的平均意见分(MOS)。本文致力于提升MOS预测模型的性能,提出一种新的评分聚合方法,解决传统MOS标注(通常为1到5分)的局限性。我们假设标注者内心考虑的是连续评分,然后选择最接近的离散等级。通过建模这一过程,我们对潜在连续分布进行量化,拟合出评分生成分布。随后,利用量化分布与实际标注之间的损失,估计该潜在分布的峰值作为新的代表性数值,替代传统的MOS。实验表明,将MOSNet的预测目标替换为此新值,可显著提升预测性能。
原文摘要 · Abstract (English)
Speech quality assessment (SQA) aims to evaluate the quality of speech samples without relying on time-consuming listener questionnaires. Recent efforts have focused on training neural-based SQA models to predict the mean opinion score (MOS) of speech samples produced by text-to-speech or voice conversion systems. This paper targets the enhancement of MOS prediction models' performance. We propose a novel score aggregation method to address the limitations of conventional annotations for MOS, which typically involve ratings on a scale from 1 to 5. Our method is based on the hypothesis that annotators internally consider continuous scores and then choose the nearest discrete rating. By modeling this process, we approximate the generative distribution of ratings by quantizing the latent continuous distribution. We then use the peak of this latent distribution, estimated through the loss between the quantized distribution and annotated ratings, as a new representative value instead of MOS. Experimental results demonstrate that substituting MOSNet's predicted target with this proposed value improves prediction performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。