arXiv:2605.03792cs.CL2026-05

首个评估司法场景中大模型风险的韩语基准,揭示其在判例检索中的严重缺陷。

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

论文配图:TriBench-Ko: Evaluating LLM Risks in Judicial Workflows
图 1 · 摘自论文原文
  • 构建四类真实司法任务的韩语评估集,覆盖判例摘要、判例检索等核心流程
  • 发现多数模型在判例检索中表现差,常遗漏关键法律信息或产生幻觉
  • 适合法律AI安全研究者、司法系统开发者参考,警示部署风险

大型语言模型(LLMs)正越来越多地融入司法工作流程。然而,现有基准主要关注模拟考试或分类任务,无法反映日常司法操作中的真实性能与风险。为此,我们公开发布TriBench-Ko,一个面向韩语司法场景的基准,用于评估大模型在实际部署中的潜在风险。该基准涵盖四个核心任务:法学推理摘要、判例检索、法律问题提取和证据分析,并综合评估多种风险类型,包括不准确(幻觉、遗漏、法条误用)、偏见(人口统计偏差、过度合规)、不一致(提示敏感性、非确定性)以及裁判越权。每项任务均基于真实司法判决设计,系统评估模型表现与特定风险。对多种主流大模型的评估显示,多数模型在判例检索中表现不佳,且难以捕捉关键法律信息。我们提供全面诊断,指出司法场景下大模型输出需严格审查的关键环节。数据集与代码已开源:https://github.com/holi-lab/TriBench-Ko

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar examination performance or classification, which fail to capture the performance and risks inherent in day-to-day judicial processes. To address this, we publicly release TriBench-Ko, a Korean benchmark designed to evaluate potential deployment risks of LLMs within the context of verified judicial task requirements. It covers four core tasks: jurisprudence summarization, precedent retrieval, legal issue extraction, and evidence analysis. It jointly assesses model behavior across multiple deployment risk categories, including inaccuracy (hallucination, omission, statutory misapplication), biases (demographic, overcompliance), inconsistencies (prompt sensitivity, non-determinism), and adjudicative overreach. Each item is structured to systematically assess both task performance and a specific risk type based on real judicial decisions. Our evaluation of a range of contemporary LLMs reveals that many models frequently manifest significant risks, most notably struggling with precedent retrieval and failing to capture critical legal information. We provide a comprehensive diagnosis of these LLMs and pinpoint critical areas where LLM-generated outputs in judicial contexts necessitate rigorous inspection and caution. Our dataset and code are available at https://github.com/holi-lab/TriBench-Ko

司法AI风险评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。