首个巴马拉语语音识别基准测试,揭示当前技术仍远未达标。
Where Are We At with Automatic Speech Recognition for the Bambara Language?
- 构建一小时专业录音的巴马拉语标准评测集,模拟理想条件。
- 顶尖模型词错误率46.76%,多款主流多语言模型超100%词错误率。
- 适合关注低资源语言语音技术、跨语言迁移的研究者参考。
本文首次为巴马拉语自动语音识别(ASR)建立标准化基准,使用一小时由马里宪法文本录制的专业语音数据。该基准在近似最优声学与语言条件下设计,用于评估37个模型,涵盖专为巴马拉语训练的系统及大型商用模型。结果显示,当前ASR性能远低于部署标准:词错误率(WER)最佳为46.76%,字符错误率(CER)最佳为13.00%;部分主流多语言模型的WER甚至超过100%。由于该数据集代表了最简化、正式的口语形式,这些结果尚未在真实场景中验证。研究提供公开基准和排行榜,以推动巴马拉语语音技术透明化发展。
原文摘要 · Abstract (English)
This paper introduces the first standardized benchmark for evaluating Automatic Speech Recognition (ASR) in the Bambara language, utilizing one hour of professionally recorded Malian constitutional text. Designed as a controlled reference set under near-optimal acoustic and linguistic conditions, the benchmark was used to evaluate 37 models, ranging from Bambara-trained systems to large-scale commercial models. Our findings reveal that current ASR performance remains significantly below deployment standards in a narrow formal domain; the top-performing system in terms of Word Error Rate (WER) achieved 46.76\% and the best Character Error Rate (CER) of 13.00\% was set by another model, while several prominent multilingual models exceeded 100\% WER. These results suggest that multilingual pre-training and model scaling alone are insufficient for underrepresented languages. Furthermore, because this dataset represents a best-case scenario of the most simplified and formal form of spoken Bambara, these figures are yet to be tested against practical, real-world settings. We provide the benchmark and an accompanying public leaderboard to facilitate transparent evaluation and future research in Bambara speech technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。