arXiv:2410.04527cs.CL2024-10EMNLP被引 36

构建多方言阿拉伯语语音数据集,助力技术普惠

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

  • 社区共建大规模多方言阿拉伯语语音数据集
  • 覆盖8种方言,含转写、性别、方言标签等标注
  • 提供强基线模型,适合方言语音研究者使用

尽管语音处理取得进展,多数世界语言和方言仍缺乏支持。这一现状加剧了技术鸿沟,阻碍了技术与社会经济包容性发展。主要原因是缺乏能推动多样化语音系统发展的数据集。本文通过开展名为Casablanca的大规模社区驱动项目,致力于缓解该问题,收集并转录多方言阿拉伯语数据集。该数据集涵盖阿尔及利亚、埃及、阿联酋、约旦、毛里塔尼亚、摩洛哥、巴勒斯坦和也门共八种方言,包含转写、性别、方言类别及代码切换标注。我们还基于Casablanca开发了若干强基线模型。Casablanca项目页面可访问:www.dlnlp.ai/speech/casablanca。

原文摘要 · Abstract (English)

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of datasets that can empower diverse speech systems. In this paper, we seek to mitigate this obstacle for a number of Arabic dialects by presenting Casablanca, a large-scale community-driven effort to collect and transcribe a multi-dialectal Arabic dataset. The dataset covers eight dialects: Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni, and includes annotations for transcription, gender, dialect, and code-switching. We also develop a number of strong baselines exploiting Casablanca. The project page for Casablanca is accessible at: www.dlnlp.ai/speech/casablanca.

语音识别多方言阿拉伯语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。