用法律数据点评估法律大模型,更贴近律师真实判分方式。
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
- 将长回答拆成独立法律信息单元,实现无参考评估
- 在自研和LegalBench数据集上超越多个基线方法
- 与专家评分更一致,提升标注者间一致性,开源数据点
法律领域的大语言模型输出评估面临独特挑战,因法律分析复杂且细微。现有评估方法或依赖昂贵的参考数据,或采用标准化评分,均难以满足法律需求。尽管LLM-as-a-Judge成为有前景的评估技术,其在法律场景中的可靠性仍取决于是否符合法律行业特有的评估流程,以及对人类法律专家而言是否可信。当前方法在此方面表现不佳,差异显著。本文旨在填补这一空白:(a) 将长篇回复分解为‘法律数据点’(LDPs),即自包含的信息单元,并提出一种反映律师实际评分方式的新型无参考评估方法;(b) 在自有数据集及开源数据集LegalBench上证明该方法优于多种基线;(c) 显示该方法与人工专家评分相关性更高,有助于提升标注者间一致性;(d) 开源部分LegalBench中使用的LDP数据,支持研究复现与该领域的进一步发展。
原文摘要 · Abstract (English)
Evaluating large language model (LLM) outputs in the legal domain presents unique challenges due to the complex and nuanced nature of legal analysis. Current evaluation approaches either depend on reference data, which is costly to produce, or use standardized assessment methods, both of which have significant limitations for legal applications. Although LLM-as-a-Judge has emerged as a promising evaluation technique, its reliability and effectiveness in legal contexts depend heavily on evaluation processes unique to the legal industry and how trustworthy the evaluation appears to the human legal expert. This is where existing evaluation methods currently fail and exhibit considerable variability. This paper aims to close the gap: a) we break down lengthy responses into 'Legal Data Points' (LDPs), self-contained units of information, and introduce a novel, reference-free evaluation methodology that reflects how lawyers evaluate legal answers; b) we demonstrate that our method outperforms a variety of baselines on both our proprietary dataset and an open-source dataset (LegalBench); c) we show how our method correlates more closely with human expert evaluations and helps improve inter-annotator agreement; and finally d) we open source our Legal Data Points for a subset of LegalBench used in our experiments, allowing the research community to replicate our results and advance research in this vital area of LLM evaluation on legal question-answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。