arXiv:2506.01776cs.CLcs.AI2025-06ACL被引 7

构建多语言指令跟随评估基准,覆盖23种语言1667个任务。

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

  • 设计跨语言指令遵循评估框架,融合规则与模型双评估方式。
  • 在23种语言上测试1667个可验证指令任务,建立主流大模型基线。
  • 适合关注多语言模型评测与国际通用性研究的开发者与学者。

随着大语言模型在自然语言处理中的快速应用,指令遵循能力已成为衡量其实际价值的关键指标。然而,现有评估方法多聚焦单语言场景,忽视了多语言与跨语言环境下的挑战与差异。为此,我们提出MaXIFE:一个涵盖23种语言、包含1667个可验证指令任务的综合性评估基准,用于评估跨语言指令遵循能力。MaXIFE结合规则驱动与模型驱动评估,兼顾效率与准确性。我们利用MaXIFE对多个主流商用大语言模型进行评估,建立了未来比较的基准结果。通过提供标准化的多语言指令遵循评估工具,MaXIFE旨在推动自然语言处理领域的研究与发展。

原文摘要 · Abstract (English)

With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and differences present in multilingual and cross-lingual contexts. To address this gap, we introduce MaXIFE: a comprehensive evaluation benchmark designed to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. MaXIFE integrates both Rule-Based Evaluation and Model-Based Evaluation, ensuring a balance of efficiency and accuracy. We applied MaXIFE to evaluate several leading commercial LLMs, establishing baseline results for future comparisons. By providing a standardized tool for multilingual instruction-following evaluation, MaXIFE aims to advance research and development in natural language processing.

指令遵循多语言模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。