arXiv:2505.12864cs.CLcs.AI2025-05中稿 · ICLR被引 51

构建340份法律考试数据集,评测大模型法律推理能力。

LEXam: Benchmarking Legal Reasoning on 340 Law Exams

  • 基于116门法学院课程的340份考题,含中英文7537道题。
  • 大模型在多步结构化法律推理上表现差,尤其开放题难答。
  • 用专家验证的评分框架,可精准评估模型推理质量。

尽管测试时扩展取得进展,长文本法律推理仍是大语言模型的关键挑战。为此,我们推出LEXam,一个源自340份法律考试的新基准,覆盖116门法学院课程,涵盖多个学科与学位层次。数据集包含7,537道英德双语考题,既有开放式长文本题,也有选项数不同的选择题。开放题附有明确的解题指引,如事实识别、规则回忆或规则应用。对开放题和选择题的评估显示,当前大模型面临显著挑战,尤其在需要多步结构化推理的开放题上表现不佳。结果还表明该数据集能有效区分不同模型的能力。通过部署集成式‘模型作为裁判’范式并经人类专家严格验证,我们证明模型推理步骤可被一致且准确地评估,与人类专家判断高度一致。该评估框架提供了一种超越简单准确率的规模化法律推理质量评估方法。

原文摘要 · Abstract (English)

Long-form legal reasoning remains a key challenge for large language models (LLMs) in spite of recent advances in test-time scaling. To address this, we introduce LEXam, a novel benchmark derived from 340 law exams spanning 116 law school courses across a range of subjects and degree levels. The dataset comprises 7,537 law exam questions in English and German. It includes both long-form, open-ended questions and multiple-choice questions with varying numbers of options. Besides reference answers, the open questions are also accompanied by explicit guidance outlining the expected legal reasoning approach such as issue spotting, rule recall, or rule application. Our evaluation on both open-ended and multiple-choice questions present significant challenges for current LLMs; in particular, they notably struggle with open questions that require structured, multi-step legal reasoning. Moreover, our results underscore the effectiveness of the dataset in differentiating between models with varying capabilities. Deploying an ensemble LLM-as-a-Judge paradigm with rigorous human expert validation, we demonstrate how model-generated reasoning steps can be evaluated consistently and accurately, closely aligning with human expert assessments. Our evaluation setup provides a scalable method to assess legal reasoning quality beyond simple accuracy metrics. Project page: https://lexam-benchmark.github.io/.

法律推理大模型评测多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。