arXiv:2411.13424cs.SDcs.CL2024-11被引 4

首个阿拉伯语方言与法英混用的语音数据集,助力多语言语音识别研究。

CAFE A Novel Code switching Dataset for Algerian Dialect French and English

  • 构建阿尔及利亚方言与法英混用的自发对话数据集,覆盖多地社会语境。
  • 包含约37小时语音,小样本集2.6小时带人工标注,支持代码切换等复杂现象分析。
  • 验证主流语音模型在混用语境下的局限性,并提出有效处理方案提升识别率。

本文发布并公开了CAFE——首个针对阿尔及利亚方言与法语、英语混用的语音数据集(数据下载链接将在接收后提供)。该数据集具有三大独特性:(a) 以自然真实的人类对话风格记录,捕捉到代码切换和重叠语音等现象;(b) 面向北非阿拉伯语方言的独特语言挑战;(c) 覆盖阿尔及利亚不同地区方言差异及多元社会语言背景。总数据量约37小时,其中子集CAFE-small(2小时36分钟)附有人工标注,包括语音分段、转录、代码切换点、重叠语音及其他事件(如噪音、笑声等)。其余约34.58小时数据为伪标签转录。此外,论文评估了Whisper large-v2、3及PromptingWhisper等先进自动语音识别(ASR)模型在此类内容上的表现,并通过优化数据处理流程与解码技术,实现混合错误率(MER)0.310、字符错误率(CER)0.329、词错误率(WER)0.538的性能提升。

原文摘要 · Abstract (English)

The paper introduces and publicly releases (Data download link available after acceptance) CAFE -- the first Code-switching dataset between Algerian dialect, French, and english languages. The CAFE speech data is unique for (a) its spontaneous speaking style in vivo human-human conversation capturing phenomena like code-switching and overlapping speech, (b) addresses distinct linguistic challenges in North African Arabic dialect; (c) the CAFE captures dialectal variations from various parts of Algeria within different sociolinguistic contexts. CAFE data contains approximately 37 hours of speech, with a subset, CAFE-small, of 2 hours and 36 minutes released with manual human annotation including speech segmentation, transcription, explicit annotation of code-switching points, overlapping speech, and other events such as noises, and laughter among others. The rest approximately 34.58 hours contain pseudo label transcriptions. In addition to the data release, the paper also highlighted the challenges of using state-of-the-art Automatic Speech Recognition (ASR) models such as Whisper large-v2,3 and PromptingWhisper to handle such content. Following, we benchmark CAFE data with the aforementioned Whisper models and show how well-designed data processing pipelines and advanced decoding techniques can improve the ASR performance in terms of Mixed Error Rate (MER) of 0.310, Character Error Rate (CER) of 0.329 and Word Error Rate (WER) of 0.538.

语音识别多语言代码切换数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。