arXiv:2605.29738cs.CLcs.AI2026-05被引 2

首个跨法域法律大模型评测基准,覆盖6国165万判决文书。

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

  • 构建六国法律任务矩阵,统一评估法院类型、判决形式等五类任务。
  • 发现少样本提升不均,语言相近性无法预测迁移效果。
  • 强调标签对齐比语言族裔更影响跨语言迁移,适合法律AI研究者。

现有法律自然语言处理评测多局限于单一语言或混合差异巨大的任务,难以实现跨语境比较。本文提出Multi-Legal-Bench,首个跨法域法律评测基准,涵盖乌克兰、法国、荷兰、波兰、捷克、立陶宛六国,覆盖四大语系,基于1.65亿份完整判决文书定义五项任务:法院类型分类、判决形式分类、案件结果预测、法律规范提取与原因类别预测,映射至各国法院登记系统的结构化元数据,形成5×6任务-法域矩阵(20/30个单元格填充)。在AWS Bedrock上评估7个前沿大模型的零样本与三样本提示性能,并测试4个中小规模模型(3-12B)以分析可扩展性。结果表明:(1) 少样本收益不均衡,取决于任务-法域组合的潜力而非语言,28个模型-法域对中有8个准确率下降;(2) 无模型在任一语言中始终领先,排名随任务与法域变化;(3) 跨语言迁移不遵循语言亲缘性:乌克兰→法国(罗曼语)仅降2.0个百分点,优于乌克兰→波兰(斯拉夫语)降13.8个百分点,标签集对齐比语系更能预测迁移质量;(4) 分词器丰富度虽有2.3倍差异,但与跨语言准确率相关性极弱(r=-0.14, p=0.24),说明模型架构与预训练数据远超分词效率的影响。所有数据、提示及模型预测均已开源。

原文摘要 · Abstract (English)

Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible. We introduce Multi-Legal-Bench, the first cross-jurisdictional legal benchmark that evaluates identical tasks across six countries (Ukraine, France, Netherlands, Poland, Czech Republic, Lithuania), four language families, and 165 million full-text court decisions. The benchmark defines five tasks (court-type classification, judgment form classification, case-outcome prediction, legal norm extraction, and cause category prediction) mapped to structured metadata from national court registries, forming a deliberately sparse 5x6 task-jurisdiction matrix (20 of 30 cells filled). We evaluate 7 frontier LLMs under zero-shot and 3-shot prompting via AWS Bedrock, with 4 additional small/medium models (3-12B) for scaling analysis. Our results reveal that: (1) few-shot gains are uneven and track how much headroom a cell leaves rather than its language, with 8 of 28 judgment-form model-jurisdiction pairs losing accuracy; (2) no single model dominates any language, rankings shift with both task and jurisdiction; (3) cross-lingual few-shot transfer does not follow language proximity: UA->FR (Romance, -2.0 pp) transfers better than UA->PL (Slavic, -13.8 pp), with label-set alignment predicting transfer quality better than language family; and (4) tokenizer fertility, despite a 2.3x spread, does not significantly predict cross-lingual accuracy (r=-0.14, p=0.24), suggesting that model architecture and pretraining data dominate tokenizer efficiency. We release all data, prompts, and model predictions.

法律AI多语言评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。