arXiv:2507.13977cs.CLeess.AS2025-07中稿 · ICASSP 2025被引 1

首个兼顾现代与古典阿拉伯语的开源语音识别模型

Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic

  • 基于FastConformer架构,构建统一处理两种阿拉伯语的模型
  • 古典阿拉伯语识别准确率达新高,现代阿拉伯语表现优异
  • 首次公开可复现的训练代码与模型,推动语言技术发展

尽管阿拉伯语是使用最广泛的语言之一,但其复杂性导致阿拉伯语自动语音识别(ASR)系统发展受限,公开可用的模型数量有限。目前研究主要聚焦于现代标准阿拉伯语(MSA),对语言内部变体关注不足。本文提出一种通用的阿拉伯语语音与文本处理方法,基于该方法训练出两个新型模型:一个专用于MSA,另一个为首个公开的统一模型,同时支持MSA和古典阿拉伯语(CA)。MSA模型在相关数据集上达到最新最佳(SOTA)性能;统一模型在带符号的古典阿拉伯语识别中实现SOTA准确率,同时保持对MSA的良好表现。为促进可复现性,我们开源了模型及训练方案。

原文摘要 · Abstract (English)

Despite Arabic being one of the most widely spoken languages, the development of Arabic Automatic Speech Recognition (ASR) systems faces significant challenges due to the language's complexity, and only a limited number of public Arabic ASR models exist. While much of the focus has been on Modern Standard Arabic (MSA), there is considerably less attention given to the variations within the language. This paper introduces a universal methodology for Arabic speech and text processing designed to address unique challenges of the language. Using this methodology, we train two novel models based on the FastConformer architecture: one designed specifically for MSA and the other, the first unified public model for both MSA and Classical Arabic (CA). The MSA model sets a new benchmark with state-of-the-art (SOTA) performance on related datasets, while the unified model achieves SOTA accuracy with diacritics for CA while maintaining strong performance for MSA. To promote reproducibility, we open-source the models and their training recipes.

语音识别多语言开源模型阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。