构建非洲语音识别真实场景基准,揭示大模型在低资源语种中的能力短板。
AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition
- 引入真实环境无脚本语音数据,覆盖非洲多语言与方言。
- 按政府、金融、医疗等10个垂直领域评估模型表现,测试数字与专有名词识别。
- 首次系统评测Sahara-v2、Gemini 3 Flash等新模型在非洲场景的泛化能力。
近期大型语言模型在高资源语言上展现出强大的语音识别与翻译能力,但非洲语言在基准测试中仍严重缺失,限制了其在低资源环境中的实际应用。早期基准虽涵盖非洲语言和口音,但缺乏真实噪声环境和细粒度领域评估。我们提出AfriVox-v2,一个面向真实非洲部署条件的综合性基准。该基准为所有支持语言引入“在野”无脚本音频,并实现严格的领域垂直化,评估模型在政府、金融、医疗、农业等十类专业领域的准确率,特别针对数字和专有名词进行定向测试。我们对Sahara-v2、Gemini 3 Flash及Omnilingual CTC等新一代语音模型进行了基准测试。结果揭示了现代语音模型在特定、嘈杂的非洲语境下的真实泛化差距,为本地化语音AI开发者提供了可靠的参考蓝图。
原文摘要 · Abstract (English)
Recent large language models (LLMs) show strong speech recognition and translation capabilities for high-resource languages. However, African languages remain dramatically underrepresented in benchmarks, limiting their practical use in low-resource settings. While early benchmarks tested African languages and accents, they lacked exhaustive real-world noise and granular domain evaluations. We present AfriVox-v2, a comprehensive benchmark designed to test speech models under realistic African deployment conditions. AfriVox-v2 introduces "in the wild" unscripted audio for all supported languages. We also introduce strict domain verticalization, evaluating model accuracy across ten sectors including government, finance, health, and agriculture and conducting targeted tests on numbers and named entities. Finally, we benchmark a new generation of speech models, including Sahara-v2, Gemini 3 Flash, and the Omnilingual CTC models. Our results expose the true generalization gap of modern speech models in specialized, noisy African contexts and provide a reliable blueprint for developers building localized voice AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。