用Mamba模型高效识别低数据量的耳语与多方言语音
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
- 基于Mamba的状态空间模型,结合自监督预训练模型
- 在低耳语数据下实现耳语与正常语音的双任务高精度识别
- 适合资源有限场景下的语音识别系统开发
耳语语音识别对传统自动语音识别系统构成重大挑战,尤其在方言多样性背景下。本文提出一种高效解决方案,采用基于Mamba的状态空间模型,并结合四种微调的自监督模型(Wav2Vec2、WavLM、HuBERT、Whisper),同时应对耳语语音与方言差异问题。实验基于新加坡、美国和爱尔兰方言的耳语与正常语音数据进行训练。结果表明,该方法仅需少量耳语数据即可实现高效建模,在wTIMIT和CHAINS数据集上达到当前最佳性能。代码已公开。
原文摘要 · Abstract (English)
Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。