arXiv:2512.00333cs.CL2025-12被引 4

首个针对11种低资源印地语系语言的多选题评测基准,揭示大模型在这些语言上的表现短板。

IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages

  • 构建涵盖11种低/极低资源印地语系语言的1.3万+题目人工标注集
  • 顶级模型Gemini-2.5在该基准上平均准确率仅58%,多数模型低于45%
  • 区分知识性与语言性题目,并测试多种复杂题型适应能力

尽管大语言模型在高资源多语言任务中表现优异,但低资源及极低资源印地语系语言仍严重缺乏评估。本文提出IndicParam,一个包含超过13,000道多选题的人工标注基准,覆盖11种语言(尼泊尔语、古吉拉特语、马拉地语、奥里亚语为低资源;多格里语、迈蒂利语、拉贾斯坦语、梵语、博多语、桑塔利语、孔卡尼语为极低资源),以及梵语-英语代码混合语料。我们评估了20个大语言模型(含专有与开源模型),发现即使表现最好的 exttt{Gemini-2.5}平均准确率也仅为58%,其次为 exttt{GPT-5}(45%)和 exttt{DeepSeek-3.2}(43.1%)。我们还对每道题进行知识性或纯语言性标注,以区分事实记忆与语法能力;同时评估模型对列表匹配、论断-理由配对、序列排序等非传统题型的处理能力。该基准揭示了跨语言迁移的局限性,并为印地语系语言建立了挑战性评测标准。数据集已公开于https://huggingface.co/datasets/bharatgenai/IndicParam,评测脚本见https://github.com/ayushbits/IndicParam。

原文摘要 · Abstract (English)

While large language models excel on high-resource multilingual tasks, low- and extremely low-resource Indic languages remain severely under-evaluated. We present IndicParam, a human-curated benchmark of over 13,000 multiple-choice questions covering 11 such languages (Nepali, Gujarati, Marathi, Odia as low-resource; Dogri, Maithili, Rajasthani, Sanskrit, Bodo, Santali, Konkani as extremely low-resource) plus Sanskrit-English code-mixed set. We evaluated 20 LLMs, both proprietary and open-weights, which reveals that even the top-performing \texttt{Gemini-2.5} reaches 58\% average accuracy, followed by \texttt{GPT-5} (45) and \texttt{DeepSeek-3.2} (43.1). We additionally label each question as knowledge-oriented or purely linguistic to discriminate factual recall from grammatical proficiency. Further, we assess the ability of LLMs to handle diverse question formats-such as list-based matching, assertion-reason pairs, and sequence ordering-alongside conventional multiple-choice questions. \benchmark\ provides insights into limitations of cross-lingual transfer and establishes a challenging benchmark for Indic languages. The dataset is available at https://huggingface.co/datasets/bharatgenai/IndicParam. Scripts to run benchmark are present at https://github.com/ayushbits/IndicParam.

低资源语言评测基准大模型评估印地语系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。