构建真实法律场景的检索增强模型评测基准,推动法律AI发展。
A Reasoning-Focused Legal Retrieval Benchmark
- 设计两类贴近真实法律研究的检索任务:律师考试题与住房法规题。
- 现有检索系统在法律问答中表现仍不理想,准确率有待提升。
- 适合法律AI研发者、法学研究者及大模型评估人员参考。
随着法律界越来越多地探索大语言模型(LLMs)在各类法律应用中的潜力,法律AI开发者转向检索增强型LLM(RAG系统)以提升性能与鲁棒性。然而,专用RAG系统的发展受限于缺乏能反映法律检索与下游问答复杂性的现实基准。为此,我们提出两个新型法律RAG基准:Bar Exam QA 和 Housing Statute QA。这些任务对应真实法律研究场景,并通过模拟法律研究过程的标注流程生成。本文详述了基准构建方法及现有检索管道的表现。结果表明,法律RAG仍是极具挑战的任务,亟需进一步研究。
原文摘要 · Abstract (English)
As the legal community increasingly examines the use of large language models (LLMs) for various legal applications, legal AI developers have turned to retrieval-augmented LLMs ("RAG" systems) to improve system performance and robustness. An obstacle to the development of specialized RAG systems is the lack of realistic legal RAG benchmarks which capture the complexity of both legal retrieval and downstream legal question-answering. To address this, we introduce two novel legal RAG benchmarks: Bar Exam QA and Housing Statute QA. Our tasks correspond to real-world legal research tasks, and were produced through annotation processes which resemble legal research. We describe the construction of these benchmarks and the performance of existing retriever pipelines. Our results suggest that legal RAG remains a challenging application, thus motivating future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。