arXiv:2506.20793cs.CL2025-06ACL被引 1

用跨语言功能测试评估大模型真实表现,发现现有基准差距大。

Multi-lingual Functional Evaluation for Large Language Models

  • 将英文功能测试转译成法、西、印、阿、约鲁巴五种语言构建新基准
  • 同一模型在不同语言下性能下降15%-24%,英语和阿拉伯语最稳定
  • 揭示静态多语言评测的局限性,适合关注实际跨语言能力的研究者

大型语言模型的多语言能力通常通过静态数据集基准(如Belebele、M-MMLU、M-GSM)评估,但这些方法难以反映模型在真实多语言场景中的表现与鲁棒性。为此,我们基于现有功能测试模板,将英文基准翻译为法语、西班牙语、印地语、阿拉伯语和约鲁巴语,构建了跨语言功能基准:跨语言小学数学符号题(CL-GSM Symbolic)和跨语言指令遵循评估(CL-IFEval)。结果表明,部分静态多语言基准与功能基准表现差异显著:在英语、法语、西班牙语中,M-GSM与CL-GSM Symbolic的性能分别下降24%、17%、18%;在多个语言间,Belebele与CL-IFEval性能下降15%-24%,而M-MMLU与CL-IFEval仅下降0.5%-3%。此外,模型在不同语言间的鲁棒性差异明显,英语和阿拉伯语表现最一致。

原文摘要 · Abstract (English)

Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide an adequate understanding of the practical performance and robustness of models across multi-lingual settings. In response, we create multi-lingual functional benchmarks -- Cross-Lingual Grade School Math Symbolic (CL-GSM Symbolic) and Cross-Lingual Instruction-Following Eval (CL-IFEval)-- by translating existing functional benchmark templates from English to five additional languages that span the range of resources available for NLP: French, Spanish, Hindi, Arabic and Yoruba. Our results reveal that some static multi-lingual benchmarks capture functional performance much more closely than others (i.e. across models, there is a 24%, 17% and 18% decrease in performance between M-GSM and CL-GSM Symbolic in English, French and Spanish respectively; similarly there's a 15 - 24% performance drop across languages between Belebele and CL-IFEval, and only a 0.5% to 3% performance drop between M-MMLU and CL-IFEval). Similarly, we find that model robustness across languages varies significantly, with certain languages (eg. Arabic, English) being the most consistently well performing across evaluation iterations.

多语言评估功能测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。