arXiv:2410.13352cs.CLcs.AI2024-10被引 8

构建首个欧洲人权法院法律推理任务与数据集,评估大模型法律推断能力。

LAR-ECHR: A New Legal Argument Reasoning Task and Dataset for Cases of the European Court of Human Rights

  • 设计法律论证链选择任务,需基于案情判断下一步合理陈述。
  • 7个主流大模型在新数据集上最高仅达75.8%准确率,仍有提升空间。
  • 方法可复用于其他法系,推动跨区域法律AI评测标准化。

我们提出法律推理(LAR)任务,旨在评估大语言模型(LLMs)的法律推理能力。该任务要求根据案件事实,从多项选择中选出法律论证链条中的正确下一步陈述。我们基于欧洲人权法院(ECHR)案例构建了名为LAR-ECHR的数据集。在该数据集上评估了七个通用大模型,发现:(a)模型排名与基于美国法律的LegalBench基准一致,尽管LAR-ECHR基于欧盟法;(b)相比LegalBench,LAR-ECHR能更清晰地区分顶尖模型表现;(c)即使最佳模型GPT-4o在该任务上也仅取得75.8%准确率,表明模型仍有显著提升空间。构建LAR-ECHR的方法可推广至其他司法体系的案例。

原文摘要 · Abstract (English)

We present Legal Argument Reasoning (LAR), a novel task designed to evaluate the legal reasoning capabilities of Large Language Models (LLMs). The task requires selecting the correct next statement (from multiple choice options) in a chain of legal arguments from court proceedings, given the facts of the case. We constructed a dataset (LAR-ECHR) for this task using cases from the European Court of Human Rights (ECHR). We evaluated seven general-purpose LLMs on LAR-ECHR and found that (a) the ranking of the models is aligned with that of LegalBench, an established US-based legal reasoning benchmark, even though LAR-ECHR is based on EU law, (b) LAR-ECHR distinguishes top models more clearly, compared to LegalBench, (c) even the best model (GPT-4o) obtains 75.8% accuracy on LAR-ECHR, indicating significant potential for further model improvement. The process followed to construct LAR-ECHR can be replicated with cases from other legal systems.

法律推理大模型评测欧洲人权法数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。