构建覆盖17国方言的阿拉伯语语音识别系统,支持多语言混用场景。
Dialectal Coverage And Generalization in Arabic Speech Recognition
- 构建多变体阿拉伯语语音识别模型,涵盖11种方言及代码混用场景。
- 在17个阿拉伯国家数据上训练,性能优于现有模型。
- 开源预训练与微调模型,适合跨方言和多语言应用研究者使用。
开发鲁棒的阿拉伯语自动语音识别(ASR)系统需有效应对语言多样性挑战。现有系统主要覆盖现代标准阿拉伯语(MSA)和少数高资源方言,难以泛化至众多口语变体。阿拉伯世界不同地区普遍存在与英语、法语的代码混用现象,对单语阿拉伯语模型构成挑战。本文提出一套优化的ASR模型,可有效识别多种口语阿拉伯语变体,包括MSA、各类方言及代码混用情况。我们公开发布覆盖17个阿拉伯语国家数据的预训练模型,以及至少包含11种方言的微调版MSA与方言识别模型,还有支持代码混用中嵌入语言的多语言模型。通过在多种口语变体上评估,验证了本方法在覆盖范围与性能上的提升。
原文摘要 · Abstract (English)
Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but fall short in coverage and generalization across the multitude of spoken variants. Code-switching with English and French is also common in different regions of the Arab world, which challenges the performance of monolingual Arabic models. In this work, we introduce a suite of ASR models optimized to effectively recognize multiple variants of spoken Arabic, including MSA, various dialects, and code-switching. We provide open-source pre-trained models that cover data from 17 Arabic-speaking countries, and fine-tuned MSA and dialectal ASR models that include at least 11 variants, as well as multi-lingual ASR models covering embedded languages in code-switched utterances. We evaluate ASR performance across these spoken varieties and demonstrate both coverage and performance gains compared to prior models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。