评测大模型在阿拉伯语伊斯兰继承案中的推理能力,发现集成模型表现最佳。
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
- 用三个基座模型集成投票,提升法律推理准确性
- 最高达92.7%准确率,在挑战赛中获第三名
- 适合对宗教法律自动化感兴趣的学者与开发者
伊斯兰继承领域对穆斯林而言至关重要,关乎遗产在继承人之间的公平分配。手动计算多种情形下的份额复杂、耗时且易出错。近年来大语言模型(LLMs)的发展激发了其在复杂法律推理任务中的应用潜力。本研究评估了前沿LLMs在解读和应用伊斯兰继承法方面的推理能力。我们采用ArabicNLP QIAS 2025挑战赛提供的数据集,其中包含源自伊斯兰法律文献的阿拉伯语继承案例。评估了多种基础模型与微调模型在准确识别继承人、计算份额及依据伊斯兰法律原则进行推理方面的能力。分析显示,融合Gemini Flash 2.5、Gemini Pro 2.5和GPT o3三模型的多数投票方案,在所有难度级别上均优于其他模型,最高达到92.7%准确率,并在Qias 2025挑战赛任务1中位列第三。
原文摘要 · Abstract (English)
Islamic inheritance domain holds significant importance for Muslims to ensure fair distribution of shares between heirs. Manual calculation of shares under numerous scenarios is complex, time-consuming, and error-prone. Recent advancements in Large Language Models (LLMs) have sparked interest in their potential to assist with complex legal reasoning tasks. This study evaluates the reasoning capabilities of state-of-the-art LLMs to interpret and apply Islamic inheritance laws. We utilized the dataset proposed in the ArabicNLP QIAS 2025 challenge, which includes inheritance case scenarios given in Arabic and derived from Islamic legal sources. Various base and fine-tuned models, are assessed on their ability to accurately identify heirs, compute shares, and justify their reasoning in alignment with Islamic legal principles. Our analysis reveals that the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms all other models that we utilized across every difficulty level. It achieves up to 92.7% accuracy and secures the third place overall in Task 1 of the Qias 2025 challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。