构建首个阿拉伯语法律推理基准,评估大模型多步推断能力
ALARB: An Arabic Legal Argument Reasoning Benchmark
- 基于1.3万份沙特商业法庭案例,构建多任务法律推理数据集
- 指令微调120亿参数模型后,判决预测性能媲美GPT-4o
- 适合研究阿拉伯语法律AI、大模型推理与司法智能的学者使用
我们提出ALARB,一个用于评估大语言模型(LLMs)在阿拉伯语法律领域推理能力的数据集和任务套件。现有阿拉伯语基准虽涵盖部分知识密集型任务如检索与理解,但针对阿拉伯语大模型的多步推理、尤其在开放场景下的数据集仍严重不足。该数据集包含来自沙特阿拉伯的超过13,000份商业法院案件,每例包含事实陈述、法院推理过程、判决结果及从法规文件中提取的引用条款。我们设计了一系列挑战性任务,反映真实法律推理的复杂性,包括判决预测、多步法律论证链补全以及基于案情的事实相关法规识别。我们在代表性开源与闭源阿拉伯语LLM上进行了基准测试,并验证了该数据集在指令微调中的实用性。值得注意的是,使用ALARB对一个120亿参数模型进行指令微调后,其在判决预测与阿拉伯语判决生成上的表现显著提升,达到与GPT-4o相当的水平。
原文摘要 · Abstract (English)
We introduce ALARB, a dataset and suite of tasks designed to evaluate the reasoning capabilities of large language models (LLMs) within the Arabic legal domain. While existing Arabic benchmarks cover some knowledge-intensive tasks such as retrieval and understanding, substantial datasets focusing specifically on multistep reasoning for Arabic LLMs, especially in open-ended contexts, are lacking. The dataset comprises over 13K commercial court cases from Saudi Arabia, with each case including the facts presented, the reasoning of the court, the verdict, as well as the cited clauses extracted from the regulatory documents. We define a set of challenging tasks leveraging this dataset and reflecting the complexity of real-world legal reasoning, including verdict prediction, completion of reasoning chains in multistep legal arguments, and identification of relevant regulations based on case facts. We benchmark a representative selection of current open and closed Arabic LLMs on these tasks and demonstrate the dataset's utility for instruction tuning. Notably, we show that instruction-tuning a modest 12B parameter model using ALARB significantly enhances its performance in verdict prediction and Arabic verdict generation, reaching a level comparable to that of GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。