构建20种印度语言的权威问答基准,评估大模型本土知识水平。
L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

- 混合生成与人工校验策略,保证数据质量并实现规模化构建
- 覆盖9个领域共69,420对问答,含19种印度语言版本
- 验证了不同评测方式结果一致,适合评估多语言大模型性能
我们提出 L3Cube-IndicQuest v2,一个大规模、多语言的金标准问答基准,用于评估大语言模型在印度本土事实知识方面的能力。该基准包含3,471个基于课程的英文问答对,覆盖九个领域,数据源自教育课程、竞赛材料和专业参考书。采用结合上下文生成、LLM验证、语义去重与人工核验的混合构建策略,实现在保持标注质量前提下的可扩展数据生产。基准被翻译为19种印度语言,形成涵盖20种语言的公开数据集,共69,420个问答对。我们在三种评测协议下评估六种大模型:LLM作为裁判和两种确定性词法标准(精确子串匹配与词重叠匹配)。三种方法得出几乎相同的模型排名,表明结果不依赖于评判方式。前沿商业模型表现遥遥领先;在开源模型中,Gemma4 31B 在所有评估的印度语言上均优于专精于印度语的 Sarvam 30B。
原文摘要 · Abstract (English)
We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。