arXiv:2608.21950cs.CLcs.AI2026-08

构建多方言阿拉伯语语音数据集,支持方言与口音识别研究。

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

  • 从11国275名说话者收集多方言语音,涵盖古典与标准阿拉伯语
  • 通过双层人工验证确保录音质量,提供结构化方言标注
  • 为方言及带口音的阿拉伯语语音识别提供强基线模型

阿拉伯语自动语音识别(ASR)因语言双轨现象、地区方言差异大以及语音资源有限而面临独特挑战。现有语音数据集通常聚焦单一方言或大规模广播/网络数据,导致语言多样性与标注质量之间的权衡。我们提出BULBUL,一个来自11个阿拉伯国家275名说话者的多方言阿拉伯语语音识别数据集。BULBUL包含结构化的方言与次方言覆盖,并收录了参与者以母语方言口音朗读的古典阿拉伯语和现代标准阿拉伯语。通过两阶段人工验证确保录音质量。我们进一步对多种近期ASR系统进行了基准测试,为现代方言及带口音阿拉伯语语音识别建立了强有力的基线。

原文摘要 · Abstract (English)

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resources. Existing speech datasets often focus on single dialects or large-scale broadcast/web data, leading to trade-offs between linguistic diversity and annotation quality. We present BULBUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries. BULBUL includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling. The quality of the recordings was ensured through a two-level human verification process. We further benchmark a range of recent ASR systems, establishing strong baselines for modern dialectal and accented Arabic ASR.

语音识别阿拉伯语多方言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。