arXiv:2606.28884eess.AS2026-06

构建680小时多语言语音数据集,评测真实场景下ASR系统表现

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

论文配图:GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
图 1 · 摘自论文原文
  • 涵盖12种低资源中东东南亚语言及中日韩等语种的多维度真实语音数据
  • 在老年、儿童及方言语音上,主流模型错误率显著上升
  • 适合评估语音识别在真实世界复杂场景下的鲁棒性

尽管现代自动语音识别(ASR)系统在高资源基准上表现优异,但其性能常高估实际应用中的鲁棒性。现有评估多孤立处理特定挑战,缺乏统一的跨领域术语、年龄差异、方言、口音及低资源语言(尤其是中东与东南亚地区)的基准,这些地区覆盖超十亿未充分评估的使用者。为此,我们提出GigaSpeechBench,一个包含680小时人工标注语音的综合性多语言、多维度真实环境下的语音识别(ASR)与语音转文本(AST)基准。该基准包含五个模块:(1)12种低资源中东和东南亚语言,以及具有挑战性的日语和韩语;(2)6种中文方言;(3)6种英语口音;(4)中英文12个垂直领域的密集术语;(5)老年人与儿童语音。同时,为支持AST评估,我们提供11种语言的人工标注中英文翻译。对领先基础模型与商用API的广泛评测显示,在这些挑战性场景中性能显著下降,暴露出关键评估盲区。

原文摘要 · Abstract (English)

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in isolation, lacking a unified benchmark for domain terminology, age variation, dialects, accents, and low-resource languages, particularly across the Middle East and Southeast Asia, representing over one billion under-evaluated speakers. To address this gap, we introduce GigaSpeechBench, a comprehensive multilingual and multidimensional in-the-wild ASR & AST benchmark comprising 680 hours of human-annotated speech. It features five modules: (1) 12 low-resource Middle Eastern and Southeast Asian languages, plus challenging Japanese and Korean; (2) 6 Chinese dialects; (3) 6 English accents; (4) dense terminology across 12 vertical domains for Chinese and English; and (5) older adult and child speech. We further provide human-annotated Chinese and English translations for 11 languages to support AST evaluation. Extensive evaluations of leading foundation models and commercial APIs reveal significant performance degradation in these challenging settings, exposing critical evaluation blind spots.

语音识别多语言真实场景评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。