针对嘈杂环境下的孟加拉语语音识别,提出混合去噪注意力模型。
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
- 融合扩散去噪与上下文交叉注意力,保留语音特征并适应不同说话人。
- 在孟加拉语语音数据集上,词错误率和字符错误率显著降低。
- 适合低资源、噪声多的孟加拉语语音识别场景使用。
孟加拉语是使用最广泛的语言之一,但在当前先进的自动语音识别(ASR)研究中仍处于边缘地位,尤其是在噪声环境和说话人多样性条件下。本文提出基于Wav2Vec-BERT的混合去噪-注意力框架BanglaRobustNet,以应对这些挑战。该架构集成基于扩散模型的去噪模块,有效抑制环境噪声同时保留孟加拉语特有的音素特征;并引入上下文交叉注意力模块,通过说话人嵌入实现对性别、年龄和方言的鲁棒性。采用端到端训练,结合CTC损失、音素一致性与说话人对齐的复合目标函数。在Mozilla Common Voice Bangla及增强噪声语音数据上的评估表明,该方法显著降低了词错误率(WER)与字符错误率(CER),验证了其有效性,确立了BanglaRobustNet作为面向低资源、高噪声语言环境的鲁棒语音识别系统。
原文摘要 · Abstract (English)
Bangla, one of the most widely spoken languages, remains underrepresented in state-of-the-art automatic speech recognition (ASR) research, particularly under noisy and speaker-diverse conditions. This paper presents BanglaRobustNet, a hybrid denoising-attention framework built on Wav2Vec-BERT, designed to address these challenges. The architecture integrates a diffusion-based denoising module to suppress environmental noise while preserving Bangla-specific phonetic cues, and a contextual cross-attention module that conditions recognition on speaker embeddings for robustness across gender, age, and dialects. Trained end-to-end with a composite objective combining CTC loss, phonetic consistency, and speaker alignment, BanglaRobustNet achieves substantial reductions in word error rate (WER) and character error rate (CER) compared to Wav2Vec-BERT and Whisper baselines. Evaluations on Mozilla Common Voice Bangla and augmented noisy speech confirm the effectiveness of our approach, establishing BanglaRobustNet as a robust ASR system tailored to low-resource, noise-prone linguistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。