构建韩法语境推理新基准,检验模型对法律时效与信息完整性的理解。
CALRK-Bench: Evaluating Context-Aware Legal Reasoning in Korean Law
- 基于韩国民法判例与咨询记录构建,评估模型对法律时效性判断能力
- 主流大模型在三类语境推理任务上表现均低于预期,最高仅42%准确率
- 适合法律AI研究者测试模型真实推理能力,非简单知识记忆
法律推理不仅需要应用法律规则,还需理解规则运行的语境。现有法律基准多假设规范固定,无法捕捉法律判断变化或多规范交互场景。本文提出CALRK-Bench,一个基于韩国法律体系的上下文感知法律推理基准。该基准评估模型识别法律规范时间有效性、判断案件信息充分性以及理解法律判决变动原因的能力。数据集源自法律判例和咨询记录,并经法律专家验证。实验显示,即使最新大语言模型在三项任务中表现普遍较差,准确率最高仅42%。CALRK-Bench为评估模型真实上下文推理能力提供新压力测试,超越单纯法律知识记忆。代码已开源:https://github.com/jhCOR/CALRKBench。
原文摘要 · Abstract (English)
Legal reasoning requires not only the application of legal rules but also an understanding of the context in which those rules operate. However, existing legal benchmarks primarily evaluate rule application under the assumption of fixed norms, and thus fail to capture situations where legal judgments shift or where multiple norms interact. In this work, we propose CALRK-Bench, a context-aware legal reasoning benchmark based on the legal system in Korean. CALRK-Bench evaluates whether models can identify the temporal validity of legal norms, determine whether sufficient legal information is available for a given case, and understand the reasons behind shifts in legal judgments. The dataset is constructed from legal precedents and legal consultation records, and is validated by legal experts. Experimental results show that even recent large language models consistently exhibit low performance on these three tasks. CALRK-Bench provides a new stress test for evaluating context-aware legal reasoning rather than simple memorization of legal knowledge. Our code is available at https://github.com/jhCOR/CALRKBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。