arXiv:2606.19597cs.SDcs.AI2026-06中稿 · INTERSPEECH 2026

用偏好预测提升语音质量评估,关键在高质量数据集

PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

  • 不依赖MOS评分,通过直接比较语音信号做偏好预测
  • 在高质量偏好数据上,模型性能显著优于基线
  • 强调高质数据的重要性,适合语音评估研究者

平均意见得分(MOS)广泛用于语音质量评估,但标量评分易受评价者差异和听感测试条件影响,引入标注噪声,限制了MOS预测的可靠性。偏好预测通过让听者直接对比信号,可减少此类变异,获得更清晰的标签。本文研究无MOS的偏好预测,提出PrefSQA,包含不确定性感知的logits、损伤注意力头及基于非匹配参考的比较模块。我们使用并优化五个数据集,包括从MOS推导和低噪声模拟的集合,涵盖匹配与非匹配内容,并实验于人工偏好集,在未见数据上测试。实验显示,对MOS衍生数据仅小幅提升;而在其他数据集上则明显优于基线,凸显高质量偏好数据的价值,验证了所提方法的有效性。

原文摘要 · Abstract (English)

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.

语音评估偏好学习高质量数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。