用WavLM和数据增强提升自然语音中嗓音力度分类准确率
Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

- 首次在嗓音力度分类中使用WavLM,结合多种数据增强策略
- 引入高斯邻近软标签,减少相邻类别混淆,最高达78.2%准确率
- 适合语音识别、情感分析等需要鲁棒声学特征的研究者
嗓音力度变化(如耳语、轻声、正常、大声、喊叫)会改变发声方式与语音声学特性,降低可懂度并影响后续语音技术的鲁棒性。分类困难源于力度呈连续变化,相邻类别易混淆,且标注数据稀缺。此前的自监督学习方法(wav2vec2、HuBERT、AST)在AVID数据集上有所提升,但仍存在边界错误。本研究首次将WavLM引入嗓音力度分类,并与wav2vec2、HuBERT进行对比。为缓解数据稀缺问题,系统评估了包括混响卷积、加性噪声、时间掩码、速度扰动、带限滤波、MixUp、CutMix在内的七种增强策略。增强策略一致提升WavLM性能,绝对增益达+0.6%至+1.8%。此外提出高斯邻近软标签,通过建模力度连续性进一步减少近边界混淆。最优模型(WavLM-BASE + 逐步解冻 + 增强 + 高斯邻近软标签)在AVID上实现78.2%平均准确率,刷新当前最佳性能。
原文摘要 · Abstract (English)
The variations in vocal effort range (e.g. whisper, soft, neutral, loud, shout) alter production and speech acoustics, reducing intelligibility and limiting the robustness of any subsequent speech technology. Classification is challenging since effort lies on a continuum, adjacent categories are easily confused, and labeled data remain scarce. Prior SSL approaches with wav2vec2, HuBERT, and AST improve performance on the AVID corpus but still suffer from boundary errors. In this study, we introduce WavLM for the first time in vocal effort classification and benchmark it against wav2vec2 and HuBERT. To address data scarcity, we conduct a systematic study of augmentation strategies, covering RIR convolution, additive noise, time masking, speed perturbation, band-limiting, MixUp, and CutMix. Augmentation consistently improves WavLM, with gains ranging from +0.6% to +1.8% absolute. We further propose Gaussian-neighbor soft labels, which further reduce near-boundary confusions by modeling the vocal effort continuum. Our best system, WavLM-BASE with gradual unfreezing, augmentation, and Gaussian-neighbor soft labels, achieves 78.2% mean accuracy, establishing a new state-of-the-art on AVID.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。