arXiv:2506.18326cs.SDeess.AS2025-06中稿 · on ICASSP 2024被引 2

用最低评分片段均值提升语音质量评估模型性能

Selecting N-lowest scores for training MOS prediction models

  • 取最低的N个主观评分求平均,替代传统平均分
  • 在VCC2018和BVCC数据集上,LCC与SRCC显著提升
  • 适合语音合成与语音转换质量评测研究者

自动语音质量评估(SQA)被广泛研究,以避免耗时的问卷调查。近年来,基于神经网络的SQA模型被用于文本转语音或语音转换生成的语音样本,主要聚焦于训练均值意见分数(MOS)预测模型。语音样本各段质量可能不一致,且人类在计算MOS时更关注低质量片段尚不明确。我们假设:人类评分时更关注低质量片段,评分差异主要源于偶然忽略劣质段落。基于此假设,我们分析了VCC2018和BVCC数据集,提出更可靠的代表性指标N_low-MOS(N个最低评分的均值)。实验表明,使用N_low-MOS训练MOSNet后,相关系数(LCC)和斯皮尔曼秩相关系数(SRCC)均优于传统MOS。结果表明,N_low-MOS更能反映主观语音质量的本质特征,使MOSNet成为更优的语音转换模型比较器。

原文摘要 · Abstract (English)

The automatic speech quality assessment (SQA) has been extensively studied to predict the speech quality without time-consuming questionnaires. Recently, neural-based SQA models have been actively developed for speech samples produced by text-to-speech or voice conversion, with a primary focus on training mean opinion score (MOS) prediction models. The quality of each speech sample may not be consistent across the entire duration, and it remains unclear which segments of the speech receive the primary focus from humans when assigning subjective evaluation for MOS calculation. We hypothesize that when humans rate speech, they tend to assign more weight to low-quality speech segments, and the variance in ratings for each sample is mainly due to accidental assignment of higher scores when overlooking the poor quality speech segments. Motivated by the hypothesis, we analyze the VCC2018 and BVCC datasets. Based on the hypothesis, we propose the more reliable representative value N_low-MOS, the mean of the $N$-lowest opinion scores. Our experiments show that LCC and SRCC improve compared to regular MOS when employing N_low-MOS to MOSNet training. This result suggests that N_low-MOS is a more intrinsic representative value of subjective speech quality and makes MOSNet a better comparator of VC models.

语音质量评估MOS预测评分机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。