arXiv:2411.09116cs.CL2024-11EMNLP被引 30

构建首个跨语言多任务并行评估基准,统一测试大模型多语能力。

P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs

  • 设计并行多语言多任务评测集,覆盖基础与专项任务
  • 提供跨语言一致的样本覆盖,支持公平对比
  • 揭示模型规模、提示词等对多语性能的影响

近期大型语言模型在翻译、代码生成、推理等多语言任务中展现出各异的能力。以往评估常局限于基础自然语言处理或单一能力任务。为此,我们提出P-MMEval,一个大规模多语言多任务基准,涵盖基础与能力专项数据集,实现各数据集间一致的语言覆盖,并提供并行样本。我们在代表性多语言模型系列上开展广泛实验,比较模型与任务间的性能差异,探究多语言表现与任务类型、模型大小、语言种类及提示词的关系,并检验英语知识向其他语言迁移的有效性。研究结果为未来多语言模型研究提供重要参考。数据集已公开于https://huggingface.co/datasets/Qwen/P-MMEval。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) showcase varied multilingual capabilities across tasks like translation, code generation, and reasoning. Previous assessments often limited their scope to fundamental natural language processing (NLP) or isolated capability-specific tasks. To alleviate this drawback, we aim to present a comprehensive multilingual multitask benchmark. First, we introduce P-MMEval, a large-scale benchmark covering effective fundamental and capability-specialized datasets. Furthermore, P-MMEval delivers consistent language coverage across various datasets and provides parallel samples. Finally, we conduct extensive experiments on representative multilingual model series to compare performances across models and tasks, explore the relationship between multilingual performances and factors such as tasks, model sizes, languages, and prompts, and examine the effectiveness of knowledge transfer from English to other languages. The resulting insights are intended to offer valuable guidance for future research. The dataset is available at https://huggingface.co/datasets/Qwen/P-MMEval.

多语言评估大模型评测跨语言迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。