为法律命题生成设计评估框架与数据集,提升LLM法律推理质量
LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation
- 构建三步评估框架,融合形式有效性与实质内容
- 100个欧盟法院判决生成命题经专家标注,显示高质量命题占比高
- 框架引导的LLM评估更接近专家判断,但细节差异仍难捕捉
法律命题生成是法律推理与学理研究的核心,但在法律自然语言处理领域仍缺乏充分研究。本文基于欧洲联盟法院判决,利用大语言模型(LLMs)开展法律命题的自动生成与评估。提出LP-Eval,一个由法律专家共同设计的三步评估框架,将法律命题质量分解为形式有效性和实质内容两个维度。基于该框架,我们发布了一个包含两位专家对100个LLM生成法律命题标注的数据集。结果表明,LLMs可生成大量结构良好且高质量的命题,且来自成熟判例的命题质量高于近期案例。进一步考察以LLM作为评估者时发现,基于评分框架引导的判断比直接整体打分更贴近专家意见,但仍无法捕捉人类专家所察觉的细微差异。
原文摘要 · Abstract (English)
Legal proposition generation is central to legal reasoning and doctrinal scholarship, yet remain under-examined in Legal NLP. This paper investigates the automatic generation and evaluation of legal propositions from decisions of the Court of Justice of the European Union using large language models (LLMs). We introduce LP-Eval, a three-step evaluation rubric co-designed with legal experts that decomposes legal proposition quality into formal validity and substantive dimensions. Using this rubric, we release a dataset of two experts' annotations for 100 LLM-generated legal propositions. Our results show that LLMs can generate predominantly well-formed and high-quality propositions, while expert evaluations reveal higher quality for propositions derived from well established cases than from recent ones. We further examine LLMs as evaluators and find that rubric-guided LLM judgments align more closely with expert assessments than direct overall scoring, but remain insensitive to finer-grained distinctions captured by human experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。