提升可学习语音特征提取的稳定性,让神经前端媲美传统方法
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
- 用STFT域掩码改进SpecAugment,缓解可学习前端过拟合
- 音频扰动结合新方法使性能提升更显著
- 两种正则化协同使用,实现与传统特征持平效果
神经前端作为自动语音识别(ASR)中替代传统固定特征提取流程的方案,可直接优化声学模型。然而其性能常落后于经典方法,我们发现主要原因是更易过拟合。本文研究针对可学习特征提取前端的正则化方法:首先分析音频扰动策略,表明可学习特征在该方法上能获得更大相对提升;其次指出标准SpecAugment在该场景下的两个局限,并提出在短时傅里叶变换(STFT)域进行掩码的简单有效改进;最后,整合两种正则化策略后,成功弥合了传统特征与可学习特征之间的性能差距。
原文摘要 · Abstract (English)
Neural front-ends are an appealing alternative to traditional, fixed feature extraction pipelines for automatic speech recognition (ASR) systems since they can be directly trained to fit the acoustic model. However, their performance often falls short compared to classical methods, which we show is largely due to their increased susceptibility to overfitting. This work therefore investigates regularization methods for training ASR models with learnable feature extraction front-ends. First, we examine audio perturbation methods and show that larger relative improvements can be obtained for learnable features. Additionally, we identify two limitations in the standard use of SpecAugment for these front-ends and propose masking in the short time Fourier transform (STFT)-domain as a simple but effective modification to address these challenges. Finally, integrating both regularization approaches effectively closes the performance gap between traditional and learnable features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。