arXiv:2511.10070cs.CL2025-11被引 6

构建了覆盖20种阿拉伯语方言的大规模数据集,提升方言识别精度。

ADI-20: Arabic Dialect Identification dataset and models

  • 基于3556小时语音数据,覆盖19种方言与标准阿拉伯语
  • 仅用30%训练数据时F1分数小幅下降,模型鲁棒性强
  • 开源数据与模型,支持方言识别研究复现与扩展

我们提出ADI-20,是先前发布的ADI-17阿拉伯语方言识别(ADI)数据集的扩展。该数据集涵盖所有阿拉伯语国家的方言,包含3,556小时来自19种阿拉伯语方言及现代标准阿拉伯语(MSA)的语音数据。我们利用该数据集训练并评估多种先进方言识别系统,包括微调基于ECAPA-TDNN的预训练模型,以及结合Whisper编码器块、注意力池化层和分类全连接层的结构。研究了训练数据量大小与模型参数数量对识别性能的影响。结果表明,在仅使用原始训练数据30%的情况下,F1分数仅有小幅下降。我们已开源所收集的数据与训练好的模型,以支持本工作的复现,并推动未来在方言识别领域的研究。

原文摘要 · Abstract (English)

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern Standard Arabic (MSA). We used this dataset to train and evaluate various state-of-the-art ADI systems. We explored fine-tuning pre-trained ECAPA-TDNN-based models, as well as Whisper encoder blocks coupled with an attention pooling layer and a classification dense layer. We investigated the effect of (i) training data size and (ii) the model's number of parameters on identification performance. Our results show a small decrease in F1 score while using only 30% of the original training data. We open-source our collected data and trained models to enable the reproduction of our work, as well as support further research in ADI.

方言识别语音数据集阿拉伯语开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。