arXiv:2502.05356eess.AScs.SD2025-02中稿 · ICASSP 2025被引 24

用教师模型指导小模型,让语音质量评估更轻量高效

Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment

  • 用自监督嵌入的教师模型生成伪标签,训练小型学生模型
  • 蒸馏可缩小性能差距至一半,模型体积缩小100倍
  • 适合资源受限场景下的语音质量评估部署

本文研究了蒸馏与剪枝方法,以减小基于自监督表示的非侵入式语音质量评估模型的规模。实验基于XLS-R-SQA模型,该模型使用wav2vec 2.0 XLS-R嵌入进行语音质量评估,并在超过10万段标注语音片段的大规模均值意见分数数据集上重新训练。对于蒸馏,以该模型为教师,利用其在未标注退化语音信号上的预测生成伪标签,训练不同规模的学生模型。对于剪枝,采用数据驱动策略。尽管数据驱动剪枝在大模型中表现更优,但针对小模型,无标签数据蒸馏更为有效。蒸馏可使学生模型与真实MOS标签的相关性差距缩小至原始基线的一半,同时相比教师模型将模型规模压缩两个数量级。

原文摘要 · Abstract (English)

In this paper, we investigate distillation and pruning methods to reduce model size for non-intrusive speech quality assessment based on self-supervised representations. Our experiments build on XLS-R-SQA, a speech quality assessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on a large compilation of mean opinion score datasets, encompassing over 100,000 labeled clips. For distillation, using this model as a teacher, we generate pseudo-labels on unlabeled degraded speech signals and train student models of varying sizes. For pruning, we use a data-driven strategy. While data-driven pruning performs better at larger model sizes, distillation on unlabeled data is more effective for smaller model sizes. Distillation can halve the gap between the baseline's correlation with ground-truth MOS labels and that of the XLS-R-based teacher model, while reducing model size by two orders of magnitude compared to the teacher model.

语音质量评估模型蒸馏轻量化自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。