arXiv:2608.26434cs.CL2026-08

首个面向非洲多语言混用语音的公开评测基准,揭示现有模型在真实场景下的严重性能下降。

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

论文配图:AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition
图 1 · 摘自论文原文
  • 构建61.36小时真实语境混用语音数据集,含16种语言及切换点标注
  • 零样本测试显示最优模型平均错误率达35.93%,无一低于24%
  • 发现语言混合模式复杂多样,单一指标无法衡量混用程度

代码混用在非洲双语对话中普遍存在,但多数语音识别系统假设输入为单语,且在精心筛选的单语基准上评估。我们提出AfriSwitch,一个包含61.36小时人工转录的野外真实混用语音数据集,涵盖16种非洲语言及语言变体,附带逐片段英语片段标签、每句话的代码混用指数(CMI)和切换点数量。统计结果显示,不同非洲语言的混用行为在两个基本独立维度上差异显著:切换频率与混合平衡度。单一标量无法全面描述语言混用程度。对五个开源及商用多语言语音识别系统进行零样本测试,其词错误率远高于同语言的已发布单语基准,最佳系统平均达35.93% WER,且无一系统在任一语言上低于24%。针对非洲地区的训练数据比模型规模或名义语言覆盖范围更能预测性能。

原文摘要 · Abstract (English)

Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, released with switch-level English span tags, perutterance Code-Mixing Index (CMI), and switch-point counts. Corpus statistics show that mixing behaviour varies widely across African languages along two largely independent axes: how often speakers alternate, and how balanced the mixture is. No single scalar captures how code-switched a language is. Benchmarking five open and commercial multilingual ASR systems zero-shot yields word error rates far above published monolingual figures for the same languages, with the best system averaging 35.93% WER and no system falling below 24% on any language. Africa-targeted training, not model scale or nominal language coverage, best predicts performance.

语音识别代码混用非洲语言多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。