扩展英文指令评测为多语言,评估大模型跨语言理解能力。
M-IFEval: Multilingual Instruction-Following Evaluation
- 构建法、日、西语指令评测集,覆盖通用与语言特异性任务。
- 8个顶尖大模型在不同语言上表现差异显著,最高差达37%。
- 适合关注多语言AI公平性与全球化应用的研究者。
指令遵循是现代大语言模型的核心能力,评估该能力对理解模型性能至关重要。现有IFEval基准采用客观标准衡量大模型表现,避免了主观的AI或人工判断,但仅包含英文指令,限制了其在其他语言中的评估能力。本文提出多语言指令遵循评估(M-IFEval)基准,涵盖法语、日语和西班牙语,包含通用及语言特异性指令。将该基准应用于8个顶尖大语言模型,结果显示不同语言和指令类型下的表现差异显著,最高达37%,凸显了多语言评测在多元文化背景下的必要性。
原文摘要 · Abstract (English)
Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Following Evaluation (IFEval) benchmark from the literature does this using objective criteria, offering a measure of LLM performance without subjective AI or human judgement. However, it only includes English instructions, limiting its ability to assess LLMs in other languages. We propose the Multilingual Instruction Following Evaluation (M-IFEval) benchmark, expanding the evaluation to French, Japanese, and Spanish, with both general and language-specific instructions. Applying this benchmark to 8 state-of-the-art LLMs, we find that benchmark performance across languages and instruction types can vary widely, underscoring the importance of a multilingual benchmark for evaluating LLMs in a diverse cultural context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。