arXiv:2507.11882cs.CL2025-07被引 1

评测30种语言下大模型的指令遵循能力,发现语言资源越少表现越差。

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

  • 构建多语言指令跟随评测集,融合翻译与人工验证保证本地化
  • 低资源语言模型准确率比高资源语言低25%-35%
  • 机器翻译数据会低估模型性能7%-22%,适合跨语言研究者使用

指令遵循能力已成为评估大语言模型的重要指标。然而,现有数据集如IFEval主要为英语单语或简单机器翻译至其他语言,难以满足多语言场景需求。本文提出针对IFEval的精细化多语言扩展——Marco-Bench-MIF,覆盖30种语言,涵盖不同本地化程度。通过结合翻译与人工验证的混合流程,解决中文等语言的大小写规范、特定地区公司名等文化与语言约束问题。对20多个大模型在该基准上的全面评估显示:(1)高/低资源语言间存在25%-35%的准确率差距;(2)模型规模影响显著,性能提升45%-60%,但脚本特异性挑战依然存在;(3)机器翻译数据导致准确率被低估7%-22%。分析揭示了跨语言指令遵循中的关键词一致性保持与复合约束遵守等核心挑战。Marco-Bench-MIF 已开源:https://github.com/AIDC-AI/Marco-Bench-MIF。

原文摘要 · Abstract (English)

Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs). However, existing datasets, such as IFEval, are either predominantly monolingual and centered on English or simply machine translated to other languages, limiting their applicability in multilingual contexts. In this paper, we present an carefully-curated extension of IFEval to a localized multilingual version named Marco-Bench-MIF, covering 30 languages with varying levels of localization. Our benchmark addresses linguistic constraints (e.g., modifying capitalization requirements for Chinese) and cultural references (e.g., substituting region-specific company names in prompts) via a hybrid pipeline combining translation with verification. Through comprehensive evaluation of 20+ LLMs on our Marco-Bench-MIF, we found that: (1) 25-35% accuracy gap between high/low-resource languages, (2) model scales largely impact performance by 45-60% yet persists script-specific challenges, and (3) machine-translated data underestimates accuracy by7-22% versus localized data. Our analysis identifies challenges in multilingual instruction following, including keyword consistency preservation and compositional constraint adherence across languages. Our Marco-Bench-MIF is available at https://github.com/AIDC-AI/Marco-Bench-MIF.

多语言指令遵循评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。