构建跨语言推理评估基准,测试大模型在29种语言中的表现差异。
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
- 基于英文基准扩展出29种语言的等量题目,支持直接跨语言对比。
- 多语言模型在低资源语言上表现下降最多达24.3%,高资源语言表现更优。
- 适合关注AI公平性、多语言能力评估的研究者和开发者使用。
现有大语言模型评估基准主要集中于英语,而当前多语言任务缺乏能专门评估跨语言推理能力的平行题目。这一双重局限使得难以全面评估模型在多语言环境下的表现。为填补此空白,我们提出MMLU-ProX,一个覆盖29种语言的综合性基准,基于英文基准构建。每种语言版本包含11,829道完全相同的题目,实现直接的跨语言比较。此外,为满足高效评估需求,还提供每语言658题的精简版。为保证质量,采用多轮大模型翻译与专家评审流程,确保表达准确、术语一致及文化适切。在此基础上,系统评估了36个顶尖大语言模型,包括增强推理与多语言优化模型。结果表明,模型在高资源语言中表现良好,但在低资源语言中性能显著下降,差距最高达24.3%。MMLU-ProX旨在推动更包容的AI发展,促进全球范围内技术的公平获取。
原文摘要 · Abstract (English)
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it challenging to comprehensively assess LLMs' performance in the multilingual setting. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-linguistic comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, with gaps of up to 24.3%. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。