针对多方言阿拉伯语语音,训练出更小却更优的自监督模型。
Ara-Best-RQ: Multi Dialectal Arabic SSL
- 基于5640小时阿拉伯语语音,用Conformer架构预训练
- 方言识别达顶尖水平,参数量少于同类模型
- 专为阿拉伯语方言设计,比通用模型效果更好
我们提出Ara-BEST-RQ,一系列专为多方言阿拉伯语语音处理设计的自监督学习(SSL)模型。利用爬取的5,640小时开放许可语音数据,并结合公开数据集,我们训练了参数量达600M的基于Conformer的BEST-RQ模型。在方言识别(DID)和自动语音识别(ASR)任务上评估,模型在方言识别上达到当前最佳性能,且参数量少于竞争模型。实验表明,针对阿拉伯语方言的专项预训练显著优于在非阿拉伯语数据上训练的多语言或单语言模型。所有模型、代码及预处理数据集将公开发布,以支持可复现性与阿拉伯语语音技术的进一步研究。
原文摘要 · Abstract (English)
We present Ara-BEST-RQ, a family of self-supervised learning (SSL) models specifically designed for multi-dialectal Arabic speech processing. Leveraging 5,640 hours of crawled Creative Commons speech and combining it with publicly available datasets, we pre-train conformer-based BEST-RQ models up to 600M parameters. Our models are evaluated on dialect identification (DID) and automatic speech recognition (ASR) tasks, achieving state-of-the-art performance on the former while using fewer parameters than competing models. We demonstrate that family-targeted pre-training on Arabic dialects significantly improves downstream performance compared to multilingual or monolingual models trained on non-Arabic data. All models, code, and pre-processed datasets will be publicly released to support reproducibility and further research in Arabic speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。