构建真实法律场景的评估基准,检验大模型的细粒度法律推理能力。
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
- 基于真实法律流程设计三类任务,覆盖咨询、案例分析与文书生成。
- 包含850个问题、12500项评分条目,支持细粒度评估。
- 发现当前主流大模型在法律推理上仍有明显短板,适合法律AI研究者使用。
随着大语言模型(LLMs)在法律领域的应用日益广泛,评估其在真实法律实践中的表现变得至关重要。然而,现有法律评测多依赖简化且高度标准化的任务,难以反映真实法律实践中的模糊性、复杂性和推理要求。此外,以往评估常采用粗粒度、单维度指标,未显式评估细粒度法律推理能力。为此,我们提出PLawBench——一个面向实际法律工作的评估基准。该基准基于真实法律工作流,涵盖公共法律咨询、实务案例分析和法律文书生成三类任务,评估模型识别法律问题与关键事实、进行结构化法律推理及生成法律上连贯文档的能力。PLawBench包含850个问题,覆盖13个实际法律场景,每个问题配有专家设计的评估量表,共约12,500项评分条目,支持细粒度评估。采用与人类专家判断对齐的基于LLM的评价器,我们评估了10个前沿大模型。实验结果表明,无一模型在PLawBench上表现良好,揭示当前大模型在细粒度法律推理方面存在显著局限,为未来法律大模型的评估与研发指明重要方向。数据已开源:https://github.com/skylenage/PLawbench。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential. However, existing legal benchmarks rely on simplified and highly standardized tasks, failing to capture the ambiguity, complexity, and reasoning demands of real legal practice. Moreover, prior evaluations often adopt coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning. To address these limitations, we introduce PLawBench, a Practical Law Benchmark designed to evaluate LLMs in realistic legal practice scenarios. Grounded in real-world legal workflows, PLawBench models the core processes of legal practitioners through three task categories: public legal consultation, practical case analysis, and legal document generation. These tasks assess a model's ability to identify legal issues and key facts, perform structured legal reasoning, and generate legally coherent documents. PLawBench comprises 850 questions across 13 practical legal scenarios, with each question accompanied by expert-designed evaluation rubrics, resulting in approximately 12,500 rubric items for fine-grained assessment. Using an LLM-based evaluator aligned with human expert judgments, we evaluate 10 state-of-the-art LLMs. Experimental results show that none achieves strong performance on PLawBench, revealing substantial limitations in the fine-grained legal reasoning capabilities of current LLMs and highlighting important directions for future evaluation and development of legal LLMs. Data is available at: https://github.com/skylenage/PLawbench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。