arXiv:2603.23750cs.CL2026-03被引 5

首个系统评估大模型伊斯兰知识能力的基准,覆盖古兰经、圣训与法学三领域。

IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge

  • 构建1万道多选题,分古兰经、圣训、法学三类,覆盖核心伊斯兰知识体系。
  • 26个模型平均准确率39.8%~93.8%,古兰经题差异最大(32.4%~99.3%)。
  • 新增学派偏见检测任务,揭示模型对不同伊斯兰法学流派偏好不一。

大型语言模型越来越多地被用于获取伊斯兰知识,但目前尚无全面的基准来评估其在核心伊斯兰学科中的表现。我们提出IslamicMMLU,一个包含10,013道多项选择题的基准,涵盖三个方向:古兰经(2,013题)、圣训(4,000题)和法理学(Fiqh,4,000题)。每个方向包含多种题型,以检验模型在不同伊斯兰知识维度上的能力。该基准用于建立IslamicMMLU公开排行榜,我们初步评估了26个大模型,其在三个方向上的平均准确率在39.8%至93.8%之间(由Gemini 3 Flash达到最高)。古兰经部分表现差异最大(32.4%至99.3%),而法理学部分引入了新颖的学派(madhab)偏见检测任务,揭示出模型在不同伊斯兰法学流派间存在明显偏好差异。阿拉伯语专用模型表现参差不齐,但均低于前沿模型。评估代码与排行榜已公开。

原文摘要 · Abstract (English)

Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple-choice questions spanning three tracks: Quran (2,013 questions), Hadith (4,000 questions), and Fiqh (jurisprudence, 4,000 questions). Each track is formed of multiple types of questions to examine LLMs capabilities handling different aspects of Islamic knowledge. The benchmark is used to create the IslamicMMLU public leaderboard for evaluating LLMs, and we initially evaluate 26 LLMs, where their averaged accuracy across the three tracks varied between 39.8% to 93.8% (by Gemini 3 Flash). The Quran track shows the widest span (99.3% to 32.4%), while the Fiqh track includes a novel madhab (Islamic school of jurisprudence) bias detection task revealing variable school-of-thought preferences across models. Arabic-specific models show mixed results, but they all underperform compared to frontier models. The evaluation code and leaderboard are made publicly available.

大模型评测伊斯兰知识多选题基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。