arXiv:2606.13751cs.CL2026-06

对比商用与开源大模型在伊斯兰继承推理中的表现差异。

Which Models Perform Better in Inheritance Reasoning?

  • 统一提示策略下比较商用与开源模型的法律推理能力。
  • 商用模型在识别继承人和规则应用上更稳定,最佳模型MRE达0.989。
  • 适合关注法律文本推理与模型可靠性评估的研究者参考。

本文介绍了团队PSL参与2026年阿拉伯伊斯兰继承推理共享任务的情况。该任务评估大语言模型在需要法律解释、多步推理和精确数值计算的继承案件中的表现。我们在统一提示策略下,对比了商用与开源模型在无需任务特定调优的结构化法律推理中的有效性。结果显示,两类模型在可靠性上存在明显差距:商用模型在识别合格继承人、应用排除规则以及保持推理步骤一致性方面表现更优;而开源模型在涉及依赖性法律决策和分数份额调整的案例中表现出更高不稳定性。最佳性能由Gemini 2.5 Flash实现,其平均相对误差(MRE)为0.989。

原文摘要 · Abstract (English)

This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning. The task evaluates the ability of large language models to solve inheritance cases that require legal interpretation, multi-step reasoning, and precise numerical computation. We compare \textit{commercial} and \textit{open-source} models under a unified prompting strategy to assess their effectiveness in structured legal reasoning with minimal task-specific adaptation. \\ Our results show a clear gap in reliability between the two model families. Commercial models demonstrate stronger performance in identifying eligible heirs, applying exclusion rules, and maintaining consistency across reasoning steps. In contrast, open-source models exhibit greater instability, particularly in cases involving dependent legal decisions and fractional share adjustments. The best performance is achieved by \textit{Gemini 2.5 Flash}, with an MRE of $0.989$.

法律推理大模型继承计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。