让语音质量检测从整体评估升级到局部识别,精准定位不同失真类型。
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
- 用部分混音策略和对比损失训练模型,生成可聚类的帧级语音质量嵌入。
- 在跨域数据上,对失真类型的识别准确率提升显著,优于传统方法。
- 适合需要精细分析语音失真来源的通信系统优化与质检场景。
自动主观语音质量评估(SSQA)传统上在语句或系统级别估计语音质量。对于过去产生中等质量语音的传输或合成系统而言,这种粒度已足够;但现代系统生成高质量语音,其失真可能仅局部存在。通过合适的模型架构和正则化损失,以语句级标签训练的SSQA模型也能实现有用的局部质量预测。本文将此类模型扩展为生成按失真类型聚类的帧级嵌入。具体而言,在干净与受损语音的并行语料上采用部分混音策略,并施加对比损失以区分不同失真类型。在域内和域外数据上的实验表明,该方法提升了失真检测能力,并可通过分析嵌入聚类实现失真类型的识别。
原文摘要 · Abstract (English)
Automatic subjective speech quality assessment (SSQA) traditionally estimates speech quality on an utterance or system level. While this resolution was adequate for older transmission or synthesis systems that produced speech signals of mediocre quality, modern systems generate high-quality speech with degradations that may occur only locally. With suitable model architectures and regularization losses, SSQA models trained with utterance-level targets can also yield useful local predictions of speech quality. In this work, we extend such models to produce frame-level embeddings that cluster by degradation type. Specifically, we employ a partial mix-up strategy on a parallel corpus of clean and degraded utterances and apply a contrastive loss to distinguish between degradation types. Through experiments on both in- and out-of-domain data, we demonstrate that our approach improves degradation detection and enables the identification of degradation types by analyzing embedding clusters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。