评测大模型对印度研究生级文化知识的理解能力,发现其在音乐、法律等领域的表现仍弱。
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
- 构建涵盖21个学科的1.7万道印地语问题,来自全国性研究生入学考试。
- Gemma3-27B表现最佳,整体准确率56.4%,但音乐、乐器、法律等主题仍不足60%。
- 首次系统评估模型在匹配、因果推理、排序等复杂题型上的表现,适合研究本土化AI的学者。
大语言模型已在理解、摘要、代码生成等任务上广泛评估,但其在印度语境下研究生水平、文化相关的深层知识理解能力仍缺乏探索。现有印度基准多聚焦基础事实问答,难以评估特定于印度背景的深度学科理解。本文提出ParamBench,包含超过1.7万道印地语问题,覆盖21个不同学科,主要源自全国性研究生入学考试,涉及历史、音乐、乐器、瑜伽、文学、哲学、法律等主题。此外,我们评估了大模型在多种题型下的表现:包括列表匹配、断言-理由配对、序列排序及传统单选题。对16个以上开源大模型的评估显示,Gemma3-27B取得最高整体准确率56.4%。细粒度分析表明,即使最优模型在音乐、古典乐器和法律等主题上表现仍弱,凸显文化语境推理的持续挑战。数据集与代码已公开于https://github.com/ayushbits/ParamBench。
原文摘要 · Abstract (English)
Large language models have been widely evaluated on tasks such as comprehension, summarization, code generation, etc. However, their performance on graduate-level, culturally grounded questions in the Indian context remains largely unexplored. Existing Indian benchmarks emphasise basic fact-orientated queries that offer limited assessment of a deeper disciplinary understanding tailored to the Indian setting. In this paper, we present ParamBench, consisting of more than 17K questions in the Hindi language, comprising questionnaires from 21 diverse subjects. These questions are primarily derived from a nationwide graduate-level entrance examination covering topics such as history, music, instruments, yoga, literature, philosophy, law, etc.~ specifically for the Indian context. Additionally, we assess the ability of LLMs to handle diverse question formats - such as list-based matching, assertion-reason pairs, and sequence ordering - alongside conventional multiple-choice questions. We evaluated the performance of more than 16 open source LLMs on this benchmark, observing that Gemma3-27B attains the highest overall accuracy of 56.4\%. Furthermore, subject-wise analysis indicates that even for the best-performing LLMs, performance remains weak on topics such as music, classical instruments, and law, underscoring persistent challenges in culturally grounded reasoning. The dataset and source code is present at https://github.com/ayushbits/ParamBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。