挑选自监督模型的早期层特征,能更好预测语音质量评分
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
- 从多个自监督模型中选取不同层的特征进行对比测试
- 早期层特征在语音质量评分预测上优于或等同于最后一层
- 适合需要高效低复杂度语音评估系统的研究者
自监督学习模型如Wav2Vec2、HuBERT和WavLM在语音处理中广泛应用。这些基于Transformer的模型包含多层结构,每层捕捉不同层次的表征。尽管已有研究探讨其分层表示的效率与性能,但语音质量评估(SQA)模型仍主要依赖最后一层特征,中间层未被充分挖掘。本文系统评估多个SSL模型各层特征在预测平均意见分(MOS)中的表现,将每层特征输入轻量级回归网络进行评估。实验结果一致表明,早期层特征在预测性能上优于或等同于最后一层,显著提升传统方法及当前最优MOS预测模型的表现。该发现凸显早期层选择的优势,可实现更高性能与更低系统复杂度。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each capturing different levels of representation. While prior studies explored their layer-wise representations for efficiency and performance, speech quality assessment (SQA) models predominantly rely on last-layer features, leaving intermediate layers underexamined. In this work, we systematically evaluate different layers of multiple SSL models for predicting mean-opinion-score (MOS). Features from each layer are fed into a lightweight regression network to assess effectiveness. Our experiments consistently show early-layers features outperform or match those from the last layer, leading to significant improvements over conventional approaches and state-of-the-art MOS prediction models. These findings highlight the advantages of early-layer selection, offering enhanced performance and reduced system complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。