arXiv:2412.13788cs.CL2024-12被引 8

构建开放通用阿拉伯语语音识别基准,评估多方言模型性能。

Open Universal Arabic ASR Leaderboard

  • 建立跨多方言数据集的持续评测平台
  • 分析模型鲁棒性、适配效率与内存消耗
  • 为社区提供统一评估框架,适合研究者参考

近年来,语音识别(ASR)模型能力的提升和多方言数据集的出现,推动阿拉伯语语音识别向全方言统一方向发展。这一趋势凸显了在多种方言上评估模型性能的必要性,以揭示模型的泛化能力。本文介绍「Open Universal Arabic ASR Leaderboard」,一个针对开源通用阿拉伯语语音识别模型的持续基准项目,覆盖多个多方言数据集。我们还对模型的鲁棒性、说话人适应能力、推理效率及内存消耗进行了全面分析。该工作旨在为阿拉伯语语音识别社区提供模型性能的参考,并建立多方言阿拉伯语语音识别的通用评估框架。

原文摘要 · Abstract (English)

In recent years, the enhanced capabilities of ASR models and the emergence of multi-dialect datasets have increasingly pushed Arabic ASR model development toward an all-dialect-in-one direction. This trend highlights the need for benchmarking studies that evaluate model performance on multiple dialects, providing the community with insights into models' generalization capabilities. In this paper, we introduce Open Universal Arabic ASR Leaderboard, a continuous benchmark project for open-source general Arabic ASR models across various multi-dialect datasets. We also provide a comprehensive analysis of the model's robustness, speaker adaptation, inference efficiency, and memory consumption. This work aims to offer the Arabic ASR community a reference for models' general performance and also establish a common evaluation framework for multi-dialectal Arabic ASR models.

语音识别多方言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。