arXiv:2503.07539cs.CL2025-03NeurIPS被引 10

构建多语言指令遵循评测基准,揭示模型在不同语言资源下的表现差异

XIFBench: Evaluating Large Language Models on Multilingual Instruction Following

  • 基于558条带0-5个约束的指令,覆盖六种语言五类约束
  • 发现低资源语言模型性能显著低于高资源语言,尤其在文化相关任务中
  • 适合关注多语言AI评估、跨语言系统开发的研究者

大型语言模型在多种应用中展现出出色的指令遵循能力,但在多语言环境下的表现缺乏系统性研究,现有评估也未深入分析不同语言情境中的细粒度约束。我们提出XIFBench,一个全面的基于约束的多语言指令遵循评测基准,包含558条指令,涵盖内容、风格、情境、格式和数值五类约束,涉及六种语言(覆盖不同资源水平)。为实现可靠一致的跨语言评估,我们引入三项方法创新:文化可及性标注、约束级翻译验证,以及以英语需求为语义锚点的要求驱动评估。对多种LLM的广泛实验不仅量化了不同资源水平下的性能差距,还深入揭示了语言资源、约束类别、指令复杂度和文化特异性对多语言指令遵循的影响。代码与数据已公开于https://github.com/zhenyuli801/XIFBench。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking fine-grained constraint analysis across diverse linguistic contexts. We introduce XIFBench, a comprehensive constraint-based benchmark for evaluating multilingual instruction-following abilities of LLMs, comprising 558 instructions with 0-5 additional constraints across five categories (Content, Style, Situation, Format, and Numerical) in six languages spanning different resource levels. To support reliable and consistent cross-lingual evaluation, we implement three methodological innovations: cultural accessibility annotation, constraint-level translation validation, and requirement-based evaluation using English requirements as semantic anchors across languages. Extensive experiments with various LLMs not only quantify performance disparities across resource levels but also provide detailed insights into how language resources, constraint categories, instruction complexity, and cultural specificity influence multilingual instruction-following. Our code and data are available at https://github.com/zhenyuli801/XIFBench.

多语言指令遵循评测基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。