用自蒸馏提升语音质量评估模型的泛化能力
DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction
- 让模型同时预测MOS和各层隐藏表示的聚类标记
- 在域内和域外测试中均显著优于基线方法
- 适合需要高泛化性的语音质量评估场景
随着自监督学习(SSL)的发展,微调预训练的SSL模型用于平均意见分(MOS)预测已达到当前最佳性能。然而,在微调过程中,这些基于SSL的MOS预测模型常出现灾难性遗忘,且容易过拟合训练集,导致泛化能力差。本文提出DistilMOS,一种新方法,不仅预测MOS,还预测每个预训练SSL模型层的隐藏表示通过聚类得到的令牌标识(token IDs)。这些分层的令牌目标作为自蒸馏信号,使MOS预测模型能提取丰富的内部知识,从而提升预测准确性和泛化能力。实验表明,该方法在域内和域外评估中均显著优于标准的SSL-based MOS预测模型,验证了其有效性和实用性。
原文摘要 · Abstract (English)
With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However, during fine-tuning, these SSL-based MOS prediction models often suffer from catastrophic forgetting of the pretrained knowledge and tend to overfit the training set, resulting in poor generalization performance. In this study, we propose DistilMOS, a novel method that learns to predict not only MOS but also token IDs obtained by clustering the hidden representations of each layer in the pretrained SSL model. These layer-wise token targets serve as self-distillation signals that enables the MOS prediction model to extract rich internal knowledge from SSL models, enhancing both prediction accuracy and generalization capability. Experimental evaluations demonstrate that our method significantly outperforms standard SSL-based MOS prediction models on both in-domain and out-of-domain evaluations, verifying the effectiveness and practicality of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。