arXiv:2602.22125cs.CL2026-02

构建14种印地语系语言的指令遵循评估基准,填补多语言AI评估空白。

IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages

  • 基于规则和可验证指令,覆盖14种印地语系语言的生成评估。
  • 每语言约800个经人工验证样本,含翻译与本土内容生成两类数据。
  • 发现模型在跨语言和词汇任务上表现差,低资源语言显著落后于英语。

指令遵循评估基准长期以英语为主,忽视数亿印地语系语言使用者的需求。我们提出IndicIFEval,一个评估大模型在14种印地语系语言中约束生成能力的基准,采用自动可验证、基于规则的指令。该基准包含每语言约800个经人工验证的样本,分为两个互补子集:源自IFEval(Zhou et al., 2023)的翻译提示(IndicIFEval-Ground),经本地化适配;以及基于本土内容合成的指令(IndicIFEval-Synth)。我们对主流开源与专有模型进行了全面评估,涵盖推理与非推理模型。结果显示,尽管模型在格式约束遵守上表现良好,但在词汇和跨语言任务上严重受限——即便在高资源语言中,整体表现仍远逊于英语。我们已发布IndicIFEval及评估脚本,以推动多语言约束生成研究(http://github.com/ai4bharat/IndicIFEval)。

原文摘要 · Abstract (English)

Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrained generation of LLMs across 14 Indic languages using automatically verifiable, rule-based instructions. It comprises around 800 human-verified examples per language spread across two complementary subsets: IndicIFEval-Ground, translated prompts from IFEval (Zhou et al., 2023) carefully localized for Indic contexts, and IndicIFEval-Ground, synthetically generated instructions grounded in native Indic content. We conduct a comprehensive evaluation of major open-weight and proprietary models spanning both reasoning and non-reasoning models. While models maintain strong adherence to formatting constraints, they struggle significantly with lexical and cross-lingual tasks -- and despite progress in high-resource languages, instruction-following across the broader Indic family lags significantly behind English. We release IndicIFEval and its evaluation scripts to support progress on multilingual constrained generation (http://github.com/ai4bharat/IndicIFEval).

多语言指令遵循评估基准印地语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。