统一听众评分标准,用对比学习提升语音质量与情感识别精度
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
- 通过对比评分建模统一的听众评分尺度,避免平均化带来的偏差
- 在语音质量评估和连续语音情感识别上均显著提升预测性能
- 适合需要高鲁棒性评分模型的研究者与工业应用
语音质量评估(SQA)和连续语音情感识别(CSER)是语音技术中的两个关键任务,均依赖于听者评分。然而,由于个体差异,评分本身存在固有偏差。以往方法采用平均听者评分尺度,并在训练集中建模所有听者评分尺度,但平均过程会扭曲序数数据,引入潜在偏差。此外,虽学习多个听者评分尺度,却仅基于平均尺度进行推理,限制了效果。本文提出一种统一听者评分尺度的方法,利用对比分数准确捕捉语句间的评分关系。实验表明,该方法在SQA和CSER任务中均有效提升预测性能,验证了其有效性与鲁棒性。
原文摘要 · Abstract (English)
Speech Quality Assessment (SQA) and Continuous Speech Emotion Recognition (CSER) are two key tasks in speech technology, both relying on listener ratings. However, these ratings are inherently biased due to individual listener factors. Previous approaches have introduced a mean listener scoring scale and modeled all listener scoring scales in the training set. However, the mean listener approach is prone to distortion from averaging ordinal data, leading to potential biases. Moreover, learning multiple listener scoring scales while inferring based only on the mean listener scale limits effectiveness. In contrast, our method focuses on modeling a unified listener scoring scale, using comparison scores to correctly capture the scoring relationships between utterances. Experimental results show that our method effectively improves prediction performance in both SQA and CSER tasks, proving its effectiveness and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。