评测大模型在伊斯兰继承法上的推理能力,发现表现差异巨大。
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
- 用1000道多选题测试模型对伊斯兰继承规则的理解与计算能力。
- o3和Gemini 2.5准确率超90%,部分模型低于50%。
- 适合研究法律AI、伊斯兰法学或大模型推理能力的学者参考。
本文评估了大语言模型在伊斯兰继承法('ilm al-mawarith)中的知识与推理能力。我们使用包含1000道多选题的基准测试,涵盖多种继承场景,旨在检验模型理解继承语境及按伊斯兰教法分配份额的能力。结果显示显著性能差距:o3和Gemini 2.5准确率超过90%,而ALLaM、Fanar、LLaMA和Mistral均低于50%。这些差异反映了推理能力和领域适配性的不同。通过详细错误分析,识别出模型常见失败模式,包括误解继承情境、误用法律规则及领域知识不足。研究揭示了大模型在结构化法律推理中的局限性,并为提升其在伊斯兰法律推理中的表现提供了方向。代码已开源:https://github.com/bouchekif/inheritance_evaluation。
原文摘要 · Abstract (English)
This paper evaluates the knowledge and reasoning capabilities of Large Language Models in Islamic inheritance law, known as 'ilm al-mawarith. We assess the performance of seven LLMs using a benchmark of 1,000 multiple-choice questions covering diverse inheritance scenarios, designed to test models' ability to understand the inheritance context and compute the distribution of shares prescribed by Islamic jurisprudence. The results reveal a significant performance gap: o3 and Gemini 2.5 achieved accuracies above 90%, whereas ALLaM, Fanar, LLaMA, and Mistral scored below 50%. These disparities reflect important differences in reasoning ability and domain adaptation. We conduct a detailed error analysis to identify recurring failure patterns across models, including misunderstandings of inheritance scenarios, incorrect application of legal rules, and insufficient domain knowledge. Our findings highlight limitations in handling structured legal reasoning and suggest directions for improving performance in Islamic legal reasoning. Code: https://github.com/bouchekif/inheritance_evaluation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。