用偏好分数评估语音质量,跨场景表现更优。
Universal Preference-Score-based Pairwise Speech Quality Assessment
- 先预测语音绝对质量分,再用函数生成相对偏好分
- 在多种数据和测试场景下均优于基线模型
- 适合语音生成系统对比与质量评估研究者
为比较两个语音生成系统的性能,最有效的方法之一是估计其生成语音之间的偏好分数。本文提出一种通用的基于偏好分数的成对语音质量评估模型(UPPSQA),用于预测两段语音样本间的偏好分数,以判断哪个质量更好。该模型首先分别预测两段语音的绝对平均意见分(MOS),再通过偏好函数将其聚合为相对偏好分数。为应对偏好数据稀缺问题,我们基于MOS数据集构建了一个新的成对语音数据集。实验结果表明,无论在不同数据类型与标签条件下的训练场景,还是在域内与域外测试场景中,UPPSQA的预测精度均优于基线模型,展现出良好的通用性。
原文摘要 · Abstract (English)
To compare the performance of two speech generation systems, one of the most effective approaches is estimating the preference score between their generated speech. This paper proposes a novel universal preference-score-based pairwise speech quality assessment (UPPSQA) model, aimed at predicting the preference score between paired speech samples to determine which one has better quality. The model first predicts the absolute mean opinion score (MOS) for the two speech samples separately, and then aggregates them into a relative preference score using a preference function. To address the scarcity of preference data, we also construct a new pairwise speech dataset based on a MOS dataset for experiments. Experimental results confirm that, whether in training scenarios with different data types and label conditions, or in both in-domain and out-of-domain test scenarios, the prediction accuracy of UPP-SQA outperforms that of the baseline models, demonstrating its universality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。