语音质量模型的隐层能自动聚类不同失真类型,无需专门训练即可识别。
Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models
- 利用DNN隐层特征进行失真分类,无需额外标注
- 在16类失真上达到94%准确率,50类真实噪声为54%
- 新模型DNSMOS+验证了性能提升对分类有益
本文实验发现:基于深度神经网络(DNN)的语音质量评估(SQA)模型具有内在的隐层表示,其中多种失真类型呈现聚集现象。尽管这些SQA模型并未针对失真分类进行训练,但实验表明在合适的SQA隐层空间中仍可获得良好的分类效果。我们通过多种音频退化类型(包括不同噪声、波形截断、增益跳变、音高偏移、压缩、混响等)验证了失真聚类特性。采用标准k近邻(kNN)分类器在SQA隐层空间中进行分类可视化。此外,我们提出一种新型DNN-based SQA模型DNSMOS+,以检验SQA性能提升是否带来分类能力增强。在包含16种失真的LibriAugmented数据集上,分类准确率达94%;在包含50类真实噪声的ESC-50数据集上,准确率为54%。
原文摘要 · Abstract (English)
In this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many types of impairments are clustered. While DNN-based SQA models are not trained for impairment classification, our experiments show good impairment classification results in an appropriate SQA latent representation. We investigate the clustering of impairments using various kinds of audio degradations that include different types of noises, waveform clipping, gain transition, pitch shift, compression, reverberation, etc. To visualize the clusters we perform classification of impairments in the SQA-latent representation domain using a standard k-nearest neighbor (kNN) classifier. We also develop a new DNN-based SQA model, named DNSMOS+, to examine whether an improvement in SQA leads to an improvement in impairment classification. The classification accuracy is 94% for LibriAugmented dataset with 16 types of impairments and 54% for ESC-50 dataset with 50 types of real noises.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。