测试三款法律AI工具,发现自研工具表现最好,但多数错误实为人工遗漏。
Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys
- 用真实法律条文数据对比三款工具,评估其准确率与错误类型。
- 自研工具STARA准确率达83%,商业平台仅58%-64%,低于通用RAG。
- 发现大量所谓错误实为人工疏漏,实际准确率可达92%,适合法律研究者参考。
检索增强生成(RAG)在法律AI中潜力巨大,但系统性评测仍不足。先前工作通过美国劳工部(DOL)律师历时数月的人工整理,构建了劳动法基准数据集LaborBench,用于评测标准RAG模型,发现其在布尔任务上准确率为70%。本文首次评估了三种新兴工具:自研法定研究助手STARA、Westlaw和LexisNexis的商用AI工具。结果显示:STARA准确率提升至83%;商业平台表现不佳,分别仅58%(Westlaw AI)和64%(Lexis+ AI),甚至低于标准RAG。通过与DOL人工输出对比进行详尽误差分析,发现主要错误包括概念混淆、例外条款误读等推理错误,以及关键条文未召回的检索失败。更关键的是,许多被认定为错误的情况实为DOL自身遗漏,经修正后STARA实际准确率达92%。论文据此提出法律RAG的五条设计原则,为多司法管辖区法律研究的AI系统提供实践指导。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) offers significant potential for legal AI, yet systematic benchmarks are sparse. Prior work introduced LaborBench to benchmark RAG models based on ostensible ground truth from an exhaustive, multi-month, manual enumeration of all U.S. state unemployment insurance requirements by U.S. Department of Labor (DOL) attorneys. That prior work found poor performance of standard RAG (70% accuracy on Boolean tasks). Here, we assess three emerging tools not previously evaluated on LaborBench: the Statutory Research Assistant (STARA), a custom statutory research tool, and two commercial tools by Westlaw and LexisNexis marketing AI statutory survey capabilities. We make five main contributions. First, we show that STARA achieves substantial performance gains, boosting accuracy to 83%. Second, we show that commercial platforms fare poorly, with accuracy of 58% (Westlaw AI) and 64% (Lexis+ AI), even worse than standard RAG. Third, we conduct a comprehensive error analysis, comparing our outputs to those compiled by DOL attorneys, and document both reasoning errors, such as confusion between related legal concepts and misinterpretation of statutory exceptions, and retrieval failures, where relevant statutory provisions are not captured. Fourth, we discover that many apparent errors are actually significant omissions by DOL attorneys themselves, such that STARA's actual accuracy is 92%. Fifth, we chart the path forward for legal RAG through concrete design principles, offering actionable guidance for building AI systems capable of accurate multi-jurisdictional legal research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。