arXiv:2511.14255cs.CL2025-11中稿 · As a Conference Pa…被引 4

首个针对非洲英语口音的多领域多国语音识别评测基准

AfriSpeech-MultiBench: A Verticalized Multidomain Multicountry Benchmark Suite for African Accented English ASR

  • 构建覆盖10+国家、7大场景的非洲英语语音评测体系
  • 开源模型在自然对话中表现好,但对非标准发音和噪声敏感
  • 多模态大模型更抗口音,但命名实体识别仍不准确

近年来,谷歌NotebookLM和OpenAI语音转语音接口等语音智能进展推动了全球语音界面热潮。然而,目前尚无公开的应用特定模型评估体系应对非洲语言多样性。我们提出AfriSpeech-MultiBench,首个面向非洲英语口音的多领域多国评测套件,涵盖10+国家、7个应用领域:金融、法律、医疗、通用对话、客服、命名实体与幻觉鲁棒性。通过分析多个开源、闭源、单模态与多模态大模型语音识别系统,使用来自多个公开非洲英语语音数据集的自发与非自发对话,发现系统性差异:开源语音模型在自然语境中表现优异,但在嘈杂非母语对话中性能下降;多模态大模型更具口音鲁棒性,但领域特定命名实体识别能力弱;专有模型在清晰语音中精度高,但因国家和领域不同而波动显著。在非洲英语数据上微调的模型可实现媲美主流模型的准确率且延迟更低,部署优势明显,但幻觉问题仍普遍存在。本评测套件的发布将助力研究者与从业者为非洲场景选择适配的语音技术,推动弱势群体的包容性语音应用发展。

原文摘要 · Abstract (English)

Recent advances in speech-enabled AI, including Google's NotebookLM and OpenAI's speech-to-speech API, are driving widespread interest in voice interfaces globally. Despite this momentum, there exists no publicly available application-specific model evaluation that caters to Africa's linguistic diversity. We present AfriSpeech-MultiBench, the first domain-specific evaluation suite for over 100 African English accents across 10+ countries and seven application domains: Finance, Legal, Medical, General dialogue, Call Center, Named Entities and Hallucination Robustness. We benchmark a diverse range of open, closed, unimodal ASR and multimodal LLM-based speech recognition systems using both spontaneous and non-spontaneous speech conversation drawn from various open African accented English speech datasets. Our empirical analysis reveals systematic variation: open-source ASR models excels in spontaneous speech contexts but degrades on noisy, non-native dialogue; multimodal LLMs are more accent-robust yet struggle with domain-specific named entities; proprietary models deliver high accuracy on clean speech but vary significantly by country and domain. Models fine-tuned on African English achieve competitive accuracy with lower latency, a practical advantage for deployment, hallucinations still remain a big problem for most SOTA models. By releasing this comprehensive benchmark, we empower practitioners and researchers to select voice technologies suited to African use-cases, fostering inclusive voice applications for underserved communities.

语音识别非洲口音多域评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。