首个评估大模型伊斯兰法推理能力的基准,揭示其知识缺陷与幻觉风险。
IslamicLegalBench: Evaluating LLMs Knowledge and Reasoning of Islamic Law Across 1,200 Years of Islamic Pluralist Legal Traditions
- 构建覆盖7派法学流派、13类任务的伊斯兰法评测集
- 顶尖模型正确率仅68%,幻觉率达21%,多模型低于35%正确率
- 复杂任务依赖语义推理,但对错误前提容忍度高,存在误导性响应
随着数百万穆斯林转向GPT、Claude、DeepSeek等大模型获取宗教指导,一个关键问题浮现:这些AI能否可靠地推理伊斯兰法?我们提出IslamicLegalBench,首个跨7派伊斯兰法学流派、包含718个实例、涵盖13类复杂度任务的基准。对9个前沿模型的评估显示重大局限:最佳模型正确率仅68%,幻觉率达21%;部分模型正确率低于35%,幻觉率超55%。少样本提示仅使2个模型提升超过1%。中等复杂度任务因需精确知识而错误最多,高复杂度任务则展现语义推理能力。虚假前提检测发现6个模型在40%以上情况下接受误导性假设,表明提示工程无法弥补基础知识缺失。IslamicLegalBench为评估AI中的伊斯兰法律推理提供首个系统框架,揭示了日益依赖的工具在精神引导中的关键漏洞。
原文摘要 · Abstract (English)
As millions of Muslims turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question arises: Can these AI systems reliably reason about Islamic law? We introduce IslamicLegalBench, the first benchmark evaluating LLMs across seven schools of Islamic jurisprudence, with 718 instances covering 13 tasks of varying complexity. Evaluation of nine state-of-the-art models reveals major limitations: the best model achieves only 68% correctness with 21% hallucination, while several models fall below 35% correctness and exceed 55% hallucination. Few-shot prompting provides minimal gains, improving only 2 of 9 models by >1%. Moderate-complexity tasks requiring exact knowledge show the highest errors, whereas high-complexity tasks display apparent competence through semantic reasoning. False premise detection indicates risky sycophancy, with 6 of 9 models accepting misleading assumptions at rates above 40%. These results highlight that prompt-based methods cannot compensate for missing foundational knowledge. IslamicLegalBench offers the first systematic framework to evaluate Islamic legal reasoning in AI, revealing critical gaps in tools increasingly relied on for spiritual guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。