arXiv:2506.02627cs.CLcs.SD2025-06中稿 · Interspeech 2025被引 5

用少量标准语微调,让小模型在阿拉伯方言语音识别上表现接近大模型。

Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning

  • 用标准语数据微调小模型,显著提升方言识别效果。
  • 少量标准语数据即可使小模型性能媲美大模型,但预训练收益有限。
  • 混合方言数据训练效果与专用模型相当,适合资源匮乏场景。

尽管商业阿拉伯语语音识别系统支持现代标准阿拉伯语(MSA),但在方言识别上表现不佳。本文研究了在五种主要阿拉伯方言(海湾、黎凡特、伊拉克、埃及、马格里布)上对OpenAI Whisper进行微调的效果,使用Mozilla Common Voice作为标准语数据,MASC数据集作为方言数据。我们评估了标准语训练数据量的影响、标准语预训练的收益,以及方言专用模型与方言合并模型的对比。结果表明,少量标准语微调数据即可使小模型性能大幅提升,达到与更大非微调模型相当水平;而标准语预训练带来的提升有限,暗示标准语与方言间共享特征较少;方言合并模型的表现与方言专用模型相当。这说明在数据平衡的前提下,合并方言数据可有效缓解低资源语音识别中的数据稀缺问题,且性能损失不显著。

原文摘要 · Abstract (English)

Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-tuning OpenAI's Whisper on five major Arabic dialects (Gulf, Levantine, Iraqi, Egyptian, Maghrebi) using Mozilla Common Voice for MSA and the MASC dataset for dialectal speech. We evaluate MSA training size effects, benefits of pre-training on MSA data, and dialect-specific versus dialect-pooled models. We find that small amounts of MSA fine-tuning data yield substantial improvements for smaller models, matching larger non-fine-tuned models. While MSA pre-training shows minimal benefit, suggesting limited shared features between MSA and dialects, our dialect-pooled models perform comparably to dialect-specific ones. This indicates that pooling dialectal data, when properly balanced, can help address data scarcity in low-resource ASR without significant performance loss.

语音识别多方言数据稀缺微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。