arXiv:2508.08912cs.CLcs.AI2025-08被引 5

弱监督预训练+持续微调,突破阿拉伯语多方言语音识别瓶颈

Munsit at NADI 2025 Shared Task 2: Pushing the Boundaries of Multidialectal Arabic ASR with Weakly Supervised Pretraining and Continual Supervised Fine-tuning

  • 用1.5万小时弱标签语音预训练,覆盖标准语与多种方言
  • 混合过滤弱标签与少量高质量标注数据,实现最优性能
  • 适合低资源方言语音识别研究者参考

自动语音识别(ASR)在虚拟助手、工业自动化、客户服务和实时转录等应用中至关重要。然而,由于标注数据有限及方言多样性带来的语言复杂性,为阿拉伯语这类低资源语言开发高精度ASR系统仍面临重大挑战。本文提出一种可扩展的训练流程,结合弱监督学习与监督微调,构建鲁棒的阿拉伯语ASR模型。第一阶段在包含现代标准阿拉伯语(MSA)和多种方言(DA)的15,000小时弱标签语音上进行预训练;第二阶段采用持续监督微调,使用过滤后的弱标签数据与少量高质量标注数据混合训练。该方法在多方言阿拉伯语语音识别挑战赛中排名第一,验证了弱监督与微调结合在缓解数据稀缺问题上的有效性,为低资源、多方言语言的高质量语音识别提供了可行路径。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) plays a vital role in enabling natural human-machine interaction across applications such as virtual assistants, industrial automation, customer support, and real-time transcription. However, developing accurate ASR systems for low-resource languages like Arabic remains a significant challenge due to limited labeled data and the linguistic complexity introduced by diverse dialects. In this work, we present a scalable training pipeline that combines weakly supervised learning with supervised fine-tuning to develop a robust Arabic ASR model. In the first stage, we pretrain the model on 15,000 hours of weakly labeled speech covering both Modern Standard Arabic (MSA) and various Dialectal Arabic (DA) variants. In the subsequent stage, we perform continual supervised fine-tuning using a mixture of filtered weakly labeled data and a small, high-quality annotated dataset. Our approach achieves state-of-the-art results, ranking first in the multi-dialectal Arabic ASR challenge. These findings highlight the effectiveness of weak supervision paired with fine-tuning in overcoming data scarcity and delivering high-quality ASR for low-resource, dialect-rich languages.

语音识别多方言弱监督阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。