用对比回归提升语音质量评估模型的泛化能力。
SCOREQ: Speech Quality Assessment with Contrastive Regression
- 引入三元组损失函数,更好捕捉主观评分的连续性。
- 在多领域数据上测试,显著提升模型泛化性能。
- 适合需要跨场景稳定预测的语音质量评估任务。
本文提出 SCOREQ,一种新型无参考语音质量评估方法。该方法采用三元组损失函数进行对比回归,解决了现有先进模型在跨领域泛化上的不足。我们首先指出使用 L2 损失训练难以捕捉主观平均意见分(MOS)的连续特性;其次通过跨多个语音领域的基准评估,验证了当前主流方法缺乏泛化能力;接着阐述 SCOREQ 的设计思路,并通过渐进式实验分析架构选择的影响;最后在多种数据和领域上与先进模型对比,结果表明 SCOREQ 有效改善了语音质量预测模型的泛化性能。结论是:将三元组损失用于对比回归,不仅能提升语音质量预测的泛化能力,也适用于其他基于回归的预测任务。
原文摘要 · Abstract (English)
In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the problem of L2 loss training failing at capturing the continuous nature of the mean opinion score (MOS) labels; (ii) demonstrate the lack of generalisation through a benchmarking evaluation across several speech domains; (iii) outline our approach and explore the impact of the architectural design decisions through incremental evaluation; (iv) evaluate the final model against state of the art models for a wide variety of data and domains. The results show that the lack of generalisation observed in state of the art speech quality metrics is addressed by SCOREQ. We conclude that using a triplet loss function for contrastive regression improves generalisation for speech quality prediction models but also has potential utility across a wide range of applications using regression-based predictive models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。