针对孟加拉方言语音识别中的噪声与方言差异,提出统一优化框架。
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
- 基于WavLM,分两阶段微调:先学标准语,再增强抗噪方言识别能力。
- 在多种方言、低信噪比环境下,性能超越wav2vec 2.0和Whisper模型。
- 适合低资源高变异性语言的语音识别系统开发者参考。
孟加拉语是全球第五大使用语言,拥有超过2.7亿使用者,其自动语音识别(ASR)仍面临巨大挑战,主要受限于方言多样性与真实环境中的声学噪声。尽管自监督学习(SSL)模型在低资源语言中取得进展,但普遍缺乏预训练阶段的去噪机制及方言适配策略。本文提出一种统一框架,基于具有掩码语音去噪目标的WavLM模型,采用多阶段微调:首先在通用领域标准孟加拉语上建立语言基础,再通过针对性数据增强实现噪声鲁棒的方言识别。在涵盖多个孟加拉方言、多种模拟噪声条件(从纯净音频到低信噪比)的综合基准上评估,结果表明该框架显著优于wav2vec 2.0和大型多语言Whisper模型,确立了新基准,为其他低资源高变异语言的实用语音识别系统提供了可扩展的有效方案。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) for Bengali, the world's fifth most spoken language, remains a significant challenge, critically hindering technological accessibility for its over 270 million speakers. This challenge is compounded by two persistent and intertwined factors: the language's vast dialectal diversity and the prevalence of acoustic noise in real-world environments. While state-of-the-art self-supervised learning (SSL) models have advanced ASR for low-resource languages, they often lack explicit mechanisms to handle environmental noise during pre-training or specialized adaptation strategies for the complex phonetic and lexical variations across Bengali dialects. This paper introduces a novel, unified framework designed to address these dual challenges simultaneously. Our approach is founded on the WavLM model, which is uniquely pre-trained with a masked speech denoising objective, making it inherently robust to acoustic distortions. We propose a specialized multi-stage fine-tuning strategy that first adapts the model to general-domain standard Bengali to establish a strong linguistic foundation and subsequently specializes it for noise-robust dialectal recognition through targeted data augmentation. The framework is rigorously evaluated on a comprehensive benchmark comprising multiple Bengali dialects under a wide range of simulated noisy conditions, from clean audio to low Signal-to-Noise Ratio (SNR) levels. Experimental results demonstrate that the proposed framework significantly outperforms strong baselines, including standard fine-tuned wav2vec 2.0 and the large-scale multilingual Whisper model. This work establishes a new state-of-the-art for this task and provides a scalable, effective blueprint for developing practical ASR systems for other low-resource, high-variation languages globally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。