通过统计正则化提升语音模型在噪声下的识别能力
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
- 引入方差-不变性-协方差正则化,调整噪声语音表征统计特性
- 在LibriSpeech测试集上相对基线提升23.3%(clean)和13.2%(other)
- 适合追求高鲁棒性的语音识别系统开发者或研究者
语音基础模型(SFMs)的噪声鲁棒性仍是关键挑战,因多数模型仅在干净数据上训练,面对噪声语音时性能下降。为此,我们提出HuBERT-VIC,一种基于方差-不变性-协方差正则化(VICReg)目标的噪声鲁棒语音基础模型。该方法调整噪声语音表示的统计特性,使模型能捕捉多样化的声学特征,增强在不同噪声类型下的泛化能力。将该方法应用于HuBERT,相较于在噪声数据上预训练的基线模型,在LibriSpeech test-clean上实现23.3%的相对性能提升,在test-other上提升13.2%。
原文摘要 · Abstract (English)
Noise robustness in speech foundation models (SFMs) has been a critical challenge, as most models are primarily trained on clean data and experience performance degradation when the models are exposed to noisy speech. To address this issue, we propose HuBERT-VIC, a noise-robust SFM with variance, in-variance, and covariance regularization (VICReg) objectives. These objectives adjust the statistics of noisy speech representations, enabling the model to capture diverse acoustic characteristics and improving the generalization ability across different types of noise. When applied to HuBERT, our model shows relative performance improvements of 23.3% on LibriSpeech test-clean and 13.2% on test-other, compared to the baseline model pre-trained on noisy speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。