构建首个大规模多语言语音理解基准,覆盖102种语言的听觉分类与问答任务。
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
- 构建覆盖102种语言的语音理解数据集,含692小时对话分类与944小时问答数据
- 发现串联式语音识别+大模型系统在多语言场景下更鲁棒,闭源语音大模型表现优于传统方法
- 揭示语音识别质量与语义理解能力间的强相关性,为跨模态训练提供新思路
语音理解对全球一半无书面系统的语言至关重要。由于缺乏双模态语音-文本训练数据,低资源语言的语音识别仍不可靠。现有评估多局限于意图识别等浅层任务。为此,我们提出Fleurs-SLU,一个包含(i)102种语言的692小时话题对话分类数据,及(ii)92种语言的944小时多选问答听觉理解数据的多语言语音理解基准。我们系统评估了端到端语音分类模型、串行语音转写+大模型分类系统,以及多模态语音-大模型的表现。结果表明:串行系统在多语言场景中更具鲁棒性;经过充分预训练的语音编码器在话题分类上表现接近最优;闭源语音大模型可达到或超越串行系统性能。我们观察到稳健的多语言语音识别、有效的语音转写与强健的多语言语音理解之间存在显著相关性,提示声学与语义表示间存在协同优化潜力。
原文摘要 · Abstract (English)
Spoken language understanding (SLU) is indispensable for half of all living languages that lack a formal writing system. Unlike for high-resource languages, for these languages, we cannot offload semantic understanding of speech to the cascade of automatic speech recognition (ASR) and text-based large language models (LLMs). Even if low-resource languages possess a writing system, ASR for these languages remains unreliable due to limited bimodal speech and text training data. Nonetheless, the evaluation of multilingual SLU is limited to shallow tasks such as intent classification or language identification. This is why we present Fleurs-SLU, a multilingual SLU benchmark that encompasses (i) 692 hours of speech for topical utterance classification in 102 languages and (ii) multiple-choice question answering via listening comprehension spanning 944 hours of speech across 92 languages. We extensively evaluate end-to-end speech classification models, cascaded systems that combine speech-to-text transcription with subsequent LLM-based classification, and multimodal speech-LLMs on Fleurs-SLU. Our results show that cascaded systems are more robust in multilingual SLU, though well-pretrained speech encoders can perform competitively in topical speech classification. Closed-source speech-LLMs match or surpass the performance of cascaded systems. We observe a strong correlation between robust multilingual ASR, effective speech-to-text translation, and strong multilingual SLU, indicating mutual benefits between acoustic and semantic speech representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。