arXiv:2606.09389cs.CL2026-06中稿 · EMNLP被引 1

构建中文法律开放式任务诊断评估基准,精准定位大模型法律回答缺陷

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

论文配图:LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks
图 1 · 摘自论文原文
  • 基于12,337条专家制定的评分细则,建立六维评估框架
  • 覆盖14类法律场景,测试18个大模型在真实法律任务中的表现差异
  • 适合法律AI研究者与模型评测人员使用,助力提升法律生成可靠性

随着大语言模型在真实法律任务中应用日益广泛,评估其开放性法律回答的可靠性变得至关重要。这类任务需要上下文敏感的回应且容错率极低,因此亟需细粒度、可诊断的评估方法以识别回答质量的具体问题。我们提出LexRubric,一个面向中文开放式法律任务的基于评分标准的评估基准。该基准包含来自法律咨询和司法考试的649个实例,涵盖14类法律场景,既反映日常法律需求,也体现专业法律推理。它进一步包含12,337条由专家撰写的原子级评分标准,并统一纳入六维评估框架,实现跨任务与维度的精确评估与诊断分析。为验证评估可靠性,我们测试了多个判官模型,并将模型判断与人工判断进行对比。此外,我们在LexRubric上评估了18个近期通用及法律领域大模型。结果表明,不同模型表现出显著的能力差异,而开放式法律任务对当前大模型仍具挑战性。数据可在https://github.com/foggpoy/LexRubric获取。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks require context-sensitive answers and allow little room for error, motivating fine-grained and diagnostic evaluation that can identify specific sources of response quality failures. We introduce LexRubric, a rubric-based benchmark for evaluating open-ended Chinese legal tasks. LexRubric contains 649 instances from legal consultation and judicial examination, which reflect both everyday legal needs and professional legal reasoning and cover 14 legal scenarios. It further includes 12,337 expert-written atomic scoring criteria organized under a unified six-dimensional framework, enabling accurate evaluation and diagnostic analysis across tasks and evaluation dimensions. To validate the reliability of the evaluation, we test multiple judge models and compare model-based judgments with human judgments. We further evaluate 18 recent general and legal-domain LLMs on LexRubric. Results show that different models exhibit distinct capability profiles, and that open-ended legal tasks remain challenging for current LLMs. Data is available at: https://github.com/foggpoy/LexRubric.

法律AI评估基准大模型评测中文LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。