无需参考音频,用自监督模型同时评估语音分离质量与可懂度。
ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
- 基于自监督表示,从混合语音和分离语音中联合预测质量与可懂度。
- 在WHAMR!数据集上,WER估计误差17%,SI-SNR估计误差1.38,相关性超0.77。
- 适用于真实场景无参考音频的语音分离评估,适合语音处理研究者。
语音分离是自动语音识别(ASR)等任务的关键预处理步骤。传统评估依赖匹配的参考音频和转录文本,但无法用于无参考的真实混合语音。本文提出一种基于自监督学习(SSL)表示的无参考评估框架,利用混合语音和分离语音联合预测音质(通过尺度不变信噪比 SI-SNR)和语音可懂度(通过词错误率 WER)。在 WHAMR! 数据集上的实验显示,WER 估计的平均绝对误差(MAE)为 17%,皮尔逊相关系数(PCC)达 0.77;SI-SNR 估计的 MAE 为 1.38,PCC 达 0.95。此外,我们验证了该估计器在多种 SSL 表示下的鲁棒性。
原文摘要 · Abstract (English)
Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and corresponding transcriptions to assess audio quality and intelligibility. However, they cannot be used to evaluate real-world mixtures for which no reference exists. This paper introduces a text-free reference-free evaluation framework based on self-supervised learning (SSL) representations. The proposed framework utilize the mixture and separated tracks to predict jointly audio quality, through the Scale Invariant Signal to Noise Ratio (SI-SNR) metric, and speech intelligibility through the Word Error Rate (WER) metric. We conducted experiments on the WHAMR! dataset, which shows a WER estimation with a mean absolute error (MAE) of 17% and a Pearson correlation coefficient (PCC) of 0.77; and SI-SNR estimation with an MAE of 1.38 and PCC of 0.95. We further demonstrate the robustness of our estimator by using various SSL representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。