arXiv:2511.10090cs.CL2025-11被引 1

用微调大模型实现高精度阿拉伯语方言识别与语音识别

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

  • 用Whisper大模型加数据增强做方言识别,效果领先
  • 方言语音识别平均错误率38.54%(WER),14.53%(CER)
  • 适合关注多方言语音处理的研究者和开发者

本文介绍Elyadata与LIA联合提交的NADI 2025多方言阿拉伯语语音处理参赛方案。我们参与了口语阿拉伯语方言识别(ADI)和多方言阿拉伯语自动语音识别(ASR)两个子任务。在所有参赛者中,我们的ADI系统排名第一,多方言ASR系统排名第二。ADI系统基于Whisper-large-v3编码器进行微调并结合数据增强,在官方测试集上达到79.83%的最高准确率。对于多方言阿拉伯语ASR,我们分别对SeamlessM4T-v2 Large(埃及变体)在八个方言上进行微调,整体在测试集上获得38.54%的平均词错误率(WER)和14.53%的字符错误率(CER)。结果表明,针对特定方言进行微调的大规模预训练语音模型在阿拉伯语语音处理中具有显著有效性。

原文摘要 · Abstract (English)

This paper describes Elyadata \& LIA's joint submission to the NADI multi-dialectal Arabic Speech Processing 2025. We participated in the Spoken Arabic Dialect Identification (ADI) and multi-dialectal Arabic ASR subtasks. Our submission ranked first for the ADI subtask and second for the multi-dialectal Arabic ASR subtask among all participants. Our ADI system is a fine-tuned Whisper-large-v3 encoder with data augmentation. This system obtained the highest ADI accuracy score of \textbf{79.83\%} on the official test set. For multi-dialectal Arabic ASR, we fine-tuned SeamlessM4T-v2 Large (Egyptian variant) separately for each of the eight considered dialects. Overall, we obtained an average WER and CER of \textbf{38.54\%} and \textbf{14.53\%}, respectively, on the test set. Our results demonstrate the effectiveness of large pre-trained speech models with targeted fine-tuning for Arabic speech processing.

语音识别方言识别大模型阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。