用针对性声学增强提升小数据ASR模型鲁棒性
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
- 通过声学增强而非语言多样性提升模型泛化能力
- 在960小时数据上将未知数据词错误率降低19.24%
- 适合资源受限下构建强健语音识别基础模型的研究者
Whisper在自动语音识别(ASR)中表现出色,常归因于其680,000小时的超大规模训练集,这对多数研究者不现实。本文研究语言与声学多样性对模型鲁棒性的影响,发现语音泛化主要由声学变化驱动,而非语言丰富度。我们发现针对性声学增强方法可显著提升模型泛化能力:在960小时的Librispeech数据集上训练时,未见数据集上的词错误率(WER)最多降低19.24%。该结果表明,聚焦声学的数据增强是构建鲁棒ASR模型的可行替代方案,尤其在缺乏大规模人工语音数据时,为未来基础型ASR模型提供新路径。
原文摘要 · Abstract (English)
Whisper's robust performance in automatic speech recognition (ASR) is often attributed to its massive 680k-hour training set, an impractical scale for most researchers. In this work, we examine how linguistic and acoustic diversity in training data affect the robustness of the ASR model and reveal that transcription generalization is primarily driven by acoustic variation rather than linguistic richness. We find that targeted acoustic augmentation methods could significantly improve the generalization ability of ASR models, reducing word-error rates by up to 19.24 percent on unseen datasets when training on the 960-hour Librispeech dataset. These findings highlight strategic acoustically focused data augmentation as a promising alternative to massive datasets for building robust ASR models, offering a potential solution to future foundation ASR models when massive human speech data is lacking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。