arXiv:2502.03945cs.CL2025-02NAACL被引 15

构建非洲口音英语对话数据集,揭示语音识别在低资源场景下的性能短板

Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond

  • 构建50段模拟医疗与非医疗场景的非洲口音英语对话数据集
  • SOTA语音识别系统在非洲口音上性能下降超10%
  • 适合关注非洲语境下语音技术公平性与医疗应用的研究者

语音技术正重塑医疗、客服及机器人等领域的交互方式,但其在非洲口音英语对话上的表现仍缺乏研究。本文提出Afrispeech-Dialog,一个包含50段模拟医疗与非医疗场景的非洲口音英语对话基准数据集,用于评估自动语音识别(ASR)及相关技术。我们对长时、带口音的语音进行评测,对比原生口音表现,发现现有SOTA语音识别与说话人分离系统性能下降超过10%。同时,我们评估大型语言模型在医疗对话摘要任务中的表现,揭示了语音识别错误对下游医疗摘要质量的显著影响。该工作凸显了在低资源地区推进包容性语音人工智能的迫切需求。

原文摘要 · Abstract (English)

Speech technologies are transforming interactions across various sectors, from healthcare to call centers and robots, yet their performance on African-accented conversations remains underexplored. We introduce Afrispeech-Dialog, a benchmark dataset of 50 simulated medical and non-medical African-accented English conversations, designed to evaluate automatic speech recognition (ASR) and related technologies. We assess state-of-the-art (SOTA) speaker diarization and ASR systems on long-form, accented speech, comparing their performance with native accents and discover a 10%+ performance degradation. Additionally, we explore medical conversation summarization capabilities of large language models (LLMs) to demonstrate the impact of ASR errors on downstream medical summaries, providing insights into the challenges and opportunities for speech technologies in the Global South. Our work highlights the need for more inclusive datasets to advance conversational AI in low-resource settings.

语音识别非洲口音医疗对话数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。