构建韩语法律推理基准,分离模型推理与知识能力
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
- 设计双组件基准:选择题与开放生成题,配套判例支持
- 30+模型测试显示推理能力仍有显著差距,专业模型更优
- 适合评估法律AI推理能力,尤其关注知识无关性
我们提出韩国规范法律基准(KCL),用于在不依赖领域知识的前提下评估大语言模型的法律推理能力。KCL提供逐题对齐的判例支持,更真实地分离推理能力与参数化知识。该基准包含两部分:(1) KCL-MCQA,283道多选题,配以1,103条对齐判例;(2) KCL-Essay,169道开放生成题,配有550条对齐判例和2,739条实例级评分标准,支持自动化评估。对30多个模型的系统评估表明,尤其是KCL-Essay任务中仍存在显著差距,且专用推理模型持续优于通用模型。所有资源(数据集与评估代码)已开源至https://github.com/lbox-kr/kcl。
原文摘要 · Abstract (English)
We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KCL provides question-level supporting precedents, enabling a more faithful disentanglement of reasoning ability from parameterized knowledge. KCL consists of two components: (1) KCL-MCQA, multiple-choice problems of 283 questions with 1,103 aligned precedents, and (2) KCL-Essay, open-ended generation problems of 169 questions with 550 aligned precedents and 2,739 instance-level rubrics for automated evaluation. Our systematic evaluation of 30+ models shows large remaining gaps, particularly in KCL-Essay, and that reasoning-specialized models consistently outperform their general-purpose counterparts. We release all resources, including the benchmark dataset and evaluation code, at https://github.com/lbox-kr/kcl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。