arXiv:2505.02172cs.CL2025-05被引 3

大模型能准确识别法律判例要点,且不靠记忆。

Identifying Legal Holdings with LLMs: A Systematic Study of Performance, Scale, and Memorization

  • 用30亿到900亿参数的模型测试法律判例识别能力。
  • 最大模型达GPT4o,F1分数0.744,媲美顶尖成果。
  • 换匿名案例仍表现良好,说明非死记硬背。

随着大语言模型能力持续提升,评估其在标准基准上的表现至关重要。本研究通过一系列实验,评估了现代大模型(参数量从30亿至900亿以上)在CaseHOLD这一法律判例识别基准数据集上的表现。实验表明,模型规模越大,性能越优,如GPT4o和AmazonNovaPro分别取得0.744和0.720的宏平均F1分数,达到该数据集上最佳公开结果水平,且无需复杂训练、微调或少样本提示。为验证结果是否源于对司法意见的训练数据记忆,我们设计并使用了一种新型引用匿名化测试,在保留语义的前提下将案件名称与引用替换为虚构内容。模型在该条件下仍保持0.728的宏平均F1分数,表明其性能并非依赖记忆。这些发现揭示了大模型在法律任务中的潜力与局限,对自动化法律分析工具及法律基准的发展具有重要意义。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to advance in capabilities, it is essential to assess how they perform on established benchmarks. In this study, we present a suite of experiments to assess the performance of modern LLMs (ranging from 3B to 90B+ parameters) on CaseHOLD, a legal benchmark dataset for identifying case holdings. Our experiments demonstrate scaling effects - performance on this task improves with model size, with more capable models like GPT4o and AmazonNovaPro achieving macro F1 scores of 0.744 and 0.720 respectively. These scores are competitive with the best published results on this dataset, and do not require any technically sophisticated model training, fine-tuning or few-shot prompting. To ensure that these strong results are not due to memorization of judicial opinions contained in the training data, we develop and utilize a novel citation anonymization test that preserves semantic meaning while ensuring case names and citations are fictitious. Models maintain strong performance under these conditions (macro F1 of 0.728), suggesting the performance is not due to rote memorization. These findings demonstrate both the promise and current limitations of LLMs for legal tasks with important implications for the development and measurement of automated legal analytics and legal benchmarks.

法律AI大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。