为评估大模型韩语法律理解能力,构建包含2510个题目的韩法律基准测试。
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
- 构建涵盖7类法律知识与4类法律推理的韩语法律评测集
- 在闭卷与检索增强两种模式下测试,结果显示模型仍有提升空间
- 联合律师开发,贴近真实法律实践场景,适合法律AI研究者参考
大语言模型在法律领域表现突出,如GPT-4已通过美国统一律师资格考试。然而其在非标准化任务及非英语语言上的表现仍有限。为此,本文提出KBL基准,用于评估大模型对韩语法律语言的理解能力,包含三部分:(1) 7类法律知识任务(510个样本),(2) 4类法律推理任务(288个样本),(3) 韩国律师资格考试(4个领域,53个任务,共2510个样本)。前两部分由律师协作开发,确保符合实际法律应用场景。同时,考虑法律从业者常依赖大量法规和判例,评测在闭卷(仅靠内部知识)与检索增强生成(RAG)两种设置下进行,使用韩国法律法规与判例语料库。结果表明模型在韩语法律理解方面仍有显著改进空间。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages other than English. This underscores the need for careful evaluation of LLMs within each legal system before application. Here, we introduce KBL, a benchmark for assessing the Korean legal language understanding of LLMs, consisting of (1) 7 legal knowledge tasks (510 examples), (2) 4 legal reasoning tasks (288 examples), and (3) the Korean bar exam (4 domains, 53 tasks, 2,510 examples). First two datasets were developed in close collaboration with lawyers to evaluate LLMs in practical scenarios in a certified manner. Furthermore, considering legal practitioners' frequent use of extensive legal documents for research, we assess LLMs in both a closed book setting, where they rely solely on internal knowledge, and a retrieval-augmented generation (RAG) setting, using a corpus of Korean statutes and precedents. The results indicate substantial room and opportunities for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。