构建多语言语音文本评估基准,揭示大模型在低资源语言上的性能差距
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
- 提出mSTEB基准,覆盖语音与文本多任务多语言评估
- 高资源语言模型性能显著优于非洲及美洲/大洋洲低资源语言
- 为提升大模型多语言公平性提供实证依据,适合关注AI公平性的研究者
大型语言模型(LLMs)在多种任务中表现优异,包括语音等多模态场景。然而,其评估常局限于英语和少数高资源语言。对于低资源语言,缺乏标准化的评估基准。本文提出mSTEB,一个新基准,用于评估LLMs在语言识别、文本分类、问答和翻译等任务上,覆盖语音与文本模态的多语言表现。我们评估了Gemini 2.0 Flash、GPT-4o(Audio)以及Qwen 2 Audio、Gemma 3 27B等领先模型。结果显示,高资源语言与低资源语言之间存在显著性能差距,尤其体现在非洲及美洲/大洋洲的语言上。研究呼吁加强对这些语言的覆盖投入。
原文摘要 · Abstract (English)
Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often limited to English and a few high-resource languages. For low-resource languages, there is no standardized evaluation benchmark. In this paper, we address this gap by introducing mSTEB, a new benchmark to evaluate the performance of LLMs on a wide range of tasks covering language identification, text classification, question answering, and translation tasks on both speech and text modalities. We evaluated the performance of leading LLMs such as Gemini 2.0 Flash and GPT-4o (Audio) and state-of-the-art open models such as Qwen 2 Audio and Gemma 3 27B. Our evaluation shows a wide gap in performance between high-resource and low-resource languages, especially for languages spoken in Africa and Americas/Oceania. Our findings show that more investment is needed to address their under-representation in LLMs coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。