arXiv:2605.29170cs.CLcs.AI2026-05被引 1

首个评估大模型在乌克兰法律推理能力的基准,覆盖5类任务。

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning

  • 基于9950万判决书构建五项法律推理任务
  • 38.6个百分点提升显示少样本提示显著有效
  • 揭示准确率误导性,适合法律AI研究者参考

现有法律自然语言处理基准几乎全部以英语为中心,忽视了形态丰富、非拉丁字母语言中的模型失效问题。本文提出UA-Legal-Bench,一个针对乌克兰法律推理的五任务评估基准,数据源自全球最大的公开司法语料库之一——统一国家法院判决登记册(EDRSR),包含9950万份判决。该基准包括:(1) 案件类型分类(4类,n=2000);(2) 判决形式分类(4类,n=2000);(3) 案件结果预测(6类,n=800);(4) 法律规范提取(n=1794);(5) 原因类别预测(22类,n=1871)。我们通过AWS Bedrock对11个不同规模(3B–675B)的LLM家族进行零样本与三样本提示评估,共调用15.8万次API。结果显示,任务间少样本效果差异显著:判决定型分类提升达+38.6个百分点,但结果预测效果混合。我们发现,在不平衡任务上准确率具有误导性:最高COP准确率(62%)模型实为多数类预测器(宏平均F1仅为23%),真正表现最佳模型仅得44%宏平均F1。家族内缩放分析表明,8B模型可在表层任务上达到前沿水平,但各家族的性能跃升阈值差异极大。所有数据、提示和模型预测均已开源。

原文摘要 · Abstract (English)

Legal NLP benchmarks are overwhelmingly English-centric, leaving failure modes in morphologically rich, non-Latin-script languages undetected. We introduce UA-Legal-Bench, a five-task benchmark for evaluating large language models on Ukrainian legal reasoning, built from the Unified State Register of Court Decisions (EDRSR) -- one of the world's largest open judicial corpora (99.5 million decisions). The benchmark comprises: (1) case-type classification (4 classes, n=2,000), (2) judgment form classification (4 classes, n=2,000), (3) case-outcome prediction (6 classes, n=800), (4) legal norm extraction (n=1,794), and (5) cause category prediction (22 classes, n=1,871). We evaluate 11 LLMs (3B--675B) from five families under zero-shot and 3-shot prompting via AWS Bedrock with 158K API calls. Our results reveal sharply task-dependent few-shot effects: few-shot prompting improves judgment form classification by up to +38.6 pp but has mixed effects on outcome prediction. We show that accuracy is misleading on imbalanced legal tasks: the model with highest COP accuracy (62%) is a majority-class predictor (macro-F1: 23%), while the genuinely best model scores only 44% macro-F1. Within-family scaling analysis reveals that 8B models can match frontier performance on surface-level tasks but scaling thresholds vary dramatically across families. We release all data, prompts, and model predictions.

法律AI多语言评估基准少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。