arXiv:2603.07539cs.CL2026-03被引 9

构建首个阿拉伯语伊斯兰继承推理数据集,支持全流程法律逻辑训练与评估。

MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs

  • 构建12,500个案例的标注数据集,覆盖从继承人识别到份额计算的完整推理链
  • 提出多阶段加权评估指标MIR-E,精准捕捉推理过程中的错误传播
  • 揭示主流大模型在规则应用与场景理解上的系统性缺陷,尤其开源模型表现不足50%

伊斯兰继承法对大语言模型具有挑战性,因其需复杂、结构化的多步推理及正确应用法学规则来计算继承份额。我们提出MAWARITH,一个包含12,500个阿拉伯语继承案例的大规模标注数据集,用于训练和评估模型在完整推理链上的表现:(i) 识别合格继承人,(ii) 应用阻断(hajb)与分配规则,(iii) 计算精确继承份额。据我们所知,MAWARITH是首个专为端到端伊斯兰继承推理设计的阿拉伯语语料库与基准。不同于以往仅限选择题的数据库,MAWARITH支持全推理链,并提供基于经典法学文献与既定继承规则的逐步解答与理由说明,以及精确的份额计算。这使模型能学习生成符合现实案例的详细分步响应。为超越最终答案准确率,我们提出MIR-E(Mawarith继承推理评估),一种加权多阶段指标,可评分关键推理环节并捕捉流水线中的错误传播。我们在零样本设置下评估六种大语言模型,商业模型达到约90%,而所有开源模型均低于50%。错误分析揭示了常见失败模式:场景误读、继承人识别错误、份额分配失误,以及对awl和radd等关键规则的缺失或错误应用。MAWARITH数据集已公开于https://gitlab.com/nlpresearcher/mawarith。

原文摘要 · Abstract (English)

Islamic inheritance law is challenging for large language models because solving inheritance cases requires complex, structured, multi-step reasoning and the correct application of juristic rules to compute heirs' shares. We introduce \textit{MAWARITH}, a large-scale annotated dataset of 12,500 Arabic inheritance cases for training and evaluating models on the full reasoning chain: (i) identifying eligible heirs, (ii) applying blocking (\textit{\d{h}ajb}) and allocation rules, and (iii) computing exact inheritance shares. To the best of our knowledge, \textit{MAWARITH} is the first Arabic corpus and benchmark designed for end-to-end Islamic inheritance reasoning. Unlike prior datasets that restrict inheritance case solving to multiple-choice questions, \textit{MAWARITH} supports the full reasoning chain and provides step-by-step solutions with justifications grounded in classical juristic sources and established inheritance rules, as well as exact share calculations. This enables models to learn how to generate detailed, step-by-step responses to user queries that reflect real-world Islamic inheritance cases. To evaluate models beyond final-answer accuracy, we propose \textit{MIR-E} (Mawarith Inheritance Reasoning Evaluation), a weighted multi-stage metric that scores key reasoning stages and captures error propagation across the pipeline. We evaluate six large language models in a zero-shot setting. A commercial model achieves about 90\%, whereas all evaluated open-source models remain below 50\%. Our error analysis identifies recurring failure patterns, including scenario misinterpretation, errors in heir identification, errors in share allocation, and missing or incorrect application of key inheritance rules such as \textit{\textquotesingle awl} and \textit{radd}. The \textit{MAWARITH} dataset is publicly available at https://gitlab.com/nlpresearcher/mawarith.

法律AI伊斯兰法推理评测多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。