arXiv:2608.04703cs.CL2026-08被引 1

评测大模型对伊斯兰学术传统理解能力的多任务基准

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

  • 构建覆盖7个学科、1200年经典的多任务评估集
  • 含3465个问答,按难度与题型分层设计
  • 适合宗教认知、跨文化AI研究者使用

大型语言模型在问答、教育和研究中应用日益广泛,尤其在依赖专业典籍的传统领域。然而,在伊斯兰研究领域,关键概念、方法与学术争论所依托的核心学术传统(turath)缺乏高质量标注资源。我们提出IslamicTurathBench(ISTB),一个面向古典伊斯兰学术的多任务、多学科评估基准。由领域专家开发并审核,ISTB包含来自35部公认典籍的3,465个问答条目,涵盖超过12个世纪的学术成果,覆盖伊斯兰研究七大核心领域。为全面评估模型能力,数据集沿两个维度组织:学术需求层次(初阶、中阶、高阶)与任务形式(选择题、段落理解题、开放问答)。数据集还包含学者参考小组的汇总评分及十种系统的零样本基线表现。该基准支持在典籍、学科、难度层级与题型上可复现地评估语言模型行为,适用于历史层叠的学术领域。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as turath, lack high-quality annotated resources. We introduce IslamicTurathBench (ISTB), a multi-task, multi-discipline dataset for evaluating LLMs on classical Islamic scholarship. Developed and reviewed by domain experts, ISTB contains 3,465 question-answer items drawn from 35 recognised source works spanning more than 12 centuries of scholarship across seven key fields of Islamic Studies. To enable comprehensive profiling of model capabilities, ISTB is structured along two axes: scholarly demand (Beginner, Intermediate, and Advanced) and task format (multiple-choice questions, passage-based comprehension, and open-ended knowledge questions). ISTB includes aggregated scores from a scholarly human reference panel and zero-shot baselines from ten systems. The dataset supports reproducible evaluation of language-model behaviour across source works, disciplines, scholarly-demand levels, and question formats in a historically layered scholarly domain.

伊斯兰研究多任务评估大模型评测学术传统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。