arXiv:2502.07346cs.CL2025-02EMNLP被引 19

构建多语言大模型评估基准,测试指令遵循与代码生成等高级能力

BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models

  • 16种语言数据经三名母语者独立校对,确保质量
  • 发现模型能力在不同语言间差异显著,单纯扩容无法弥补
  • 适合关注多语言模型性能与公平性的研究者使用

现有多语言评测侧重基础理解任务,但对大语言模型(LLMs)而言,指令遵循、推理、长上下文理解、代码生成等能力更为关键。然而,跨语言衡量这些高级能力仍缺乏系统性方法。为此,我们提出BenchMAX,一个支持多语言公平比较的综合性评测基准。所有数据先由机器从英语翻译至16种其他语言,再经三位母语者独立标注,确保质量。实验显示,核心能力在不同语言间表现不一,单纯扩大模型规模无法消除差距。BenchMAX为多语言大模型发展提供重要测试平台,数据与代码已公开。

原文摘要 · Abstract (English)

Previous multilingual benchmarks focus primarily on simple understanding tasks, but for large language models(LLMs), we emphasize proficiency in instruction following, reasoning, long context understanding, code generation, and so on. However, measuring these advanced capabilities across languages is underexplored. To address the disparity, we introduce BenchMAX, a multi-way multilingual evaluation benchmark that allows for fair comparisons of these important abilities across languages. To maintain high quality, three distinct native-speaking annotators independently annotate each sample within all tasks after the data was machine-translated from English into 16 other languages. Additionally, we present a novel translation challenge stemming from dataset construction. Extensive experiments on BenchMAX reveal varying effectiveness of core capabilities across languages, highlighting performance gaps that cannot be bridged by simply scaling up model size. BenchMAX serves as a comprehensive multilingual evaluation platform, providing a promising test bed to promote the development of multilingual language models. The dataset and code are publicly accessible.

多语言评估大模型评测指令遵循跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。