用Whisper模型特征预测语音质量,效果优于现有方法。
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
- 基于Whisper语音识别模型提取特征,构建非侵入式语音质量预测器。
- 在所有NISQA测试集上与人工评分相关性更高,达0.91以上。
- 跨领域适应能力显著强于DNSMOS,适合真实场景应用。
近年来,神经网络驱动的语音质量(SQ)预测研究进展显著。尽管主要目标是开发无需参考信号的非侵入式评估指标以衡量语音增强系统性能,近期工作也探索将神经语音质量预测器直接嵌入下游语音任务的损失函数中。为支持此类预测器的训练,已构建多个包含音频及对应人工质量评分的大规模数据集。最新研究表明,由大型无监督或半监督基础语音模型生成的语音表示对神经语音质量预测任务非常有效。本文提出一种新颖且鲁棒的语音质量预测系统,基于语音识别模型(ASR)提取的特征表示,该特征被证实是语音质量预测任务中的强大输入。所提系统在所有NISQA测试集上与人工平均主观评分(MOS)的相关性均高于近期方法,且在跨域适应方面显著优于常用指标DNSMOS。
原文摘要 · Abstract (English)
There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems, recent work has also investigated the direct inference of neural SQ predictors within the loss function of downstream speech tasks. To aid in the training of SQ predictors, several large datasets of audio with corresponding human labels of quality have been created. Recent work in this area has shown that speech representations derived from large unsupervised or semi-supervised foundational speech models are useful input feature representations for neural SQ prediction. In this work, a novel and robust SQ predictor is proposed based on feature representations extracted from an ASR model, found to be a powerful input feature for the SQ prediction task. The proposed system achieves higher correlation with human MOS ratings than recent approaches on all NISQA test sets and shows significantly better domain adaption compared to the commonly used DNSMOS metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。