通过高频增强与两阶段训练,提升多采样率语音质量评估精度
SA-SSL-MOS: Self-supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment
- 采用并行分支结构融合48kHz高频信息,保留高采样率语音特征
- 在有限多速率数据下,两阶段训练使模型泛化能力显著提升
- 适合需要跨采样率语音质量评估的工业级应用
针对16-48 kHz多采样率语音的客观质量评估(SQA)任务,由于缺乏标注的多速率语音数据集,该任务极具挑战性。尽管自监督学习(SSL)已被广泛用于提升性能,但现有方法通常在16kHz语音上预训练,导致高采样率中丰富的高频信息被忽略。为此,本文提出一种基于频谱增强的自监督学习方法,通过并行分支架构引入最高达48kHz的高频特征。同时设计两阶段训练策略:先在大规模48kHz数据集上预训练,再在小规模多速率数据集上微调。实验表明,利用此前被忽视的高频信息对准确评估多速率语音质量至关重要,且该两阶段训练显著提升了模型在多速率数据受限时的泛化能力。
原文摘要 · Abstract (English)
Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. The challenge arises due to the limited availability of a MOS-labeled training dataset comprising multi-rate speech samples. While self-supervised learning (SSL) models have been widely adopted in SQA to boost performance, a key limitation is that they are pretrained on 16 kHz speech and therefore discard high-frequency information present in higher sampling rates. To address this issue, we propose a spectrogram-augmented SSL method that incorporates high-frequency features (up to 48 kHz sampling rate) through a parallel-branch architecture. We further introduce a two-step training scheme: the model is first pre-trained on a large 48 kHz dataset and then fine-tuned on a smaller multi-rate dataset. Experimental results show that leveraging high-frequency information overlooked by SSL features is crucial for accurate multi-rate SQA, and that the proposed two-step training substantially improves generalization when multi-rate data is limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。