arXiv:2605.08437cs.CLcs.AI2026-05

评测大模型在法官级法律任务中的判案能力,发现顶尖模型仍难达标。

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

论文配图:Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
图 1 · 摘自论文原文
  • 构建巴西司法考试真题库,涵盖多轮论述与完整判决书写作。
  • 23个主流大模型平均分仅6.97/10,最佳模型未达70%满分。
  • 首次以大模型为裁判,验证评估一致性高,适合法律AI研究者使用。

现有法律AI评测多聚焦于生成法律论证或文书,但评判论证——权衡对立主张、将法理适用于事实、作出合理裁决——对法律体系同样关键。我们提出Magis-Bench,一个基于2023至2025年巴西司法职位竞争性考试的基准测试,包含74道题目,涵盖具有多轮结构的论述性法律分析题及需撰写完整民事与刑事判决书的实务练习题。采用四个独立前沿大模型作为裁判,评估23个当前最先进的大模型。结果显示评委间高度一致(Kendall's W = 0.984;成对Kendall's τ ≥ 0.897),谷歌Gemini-3-Pro-Preview得分最高(6.97/10),其次为Gemini-3-Flash-Preview(6.67)和Claude-4.5-Opus(6.46)。即使最优模型也未达到满分的70%,表明当前大模型在司法级法律推理与写作方面仍具挑战性。我们公开完整基准数据、模型输出与评估代码,以推动法律AI能力研究。

原文摘要 · Abstract (English)

Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competing claims, applying doctrine to facts, and rendering reasoned decisions -- is arguably as fundamental to a well-functioning legal system as advocacy itself. We introduce Magis-Bench, a benchmark for evaluating LLMs on magistrate-level writing tasks derived from recent Brazilian competitive examinations for judicial positions. Magis-Bench comprises 74 questions from eight examinations conducted between 2023 and 2025, including discursive legal analysis questions with multi-turn structure and practical exercises requiring the composition of complete civil and criminal judicial sentences. We evaluate 23 state-of-the-art LLMs using an LLM-as-a-judge methodology with four independent frontier models as evaluators. Our results show strong inter-judge agreement (Kendall's $W = 0.984$; pairwise Kendall's $τ\ge 0.897$), with Google's Gemini-3-Pro-Preview achieving the highest average score (6.97/10), followed by Gemini-3-Flash-Preview (6.67) and Claude-4.5-Opus (6.46). Even the best-performing models score below 70\% of the maximum, indicating that judicial-level legal reasoning and writing remain challenging for current LLMs. We release the complete benchmark, model outputs, and evaluation code to support further research on legal AI capabilities.

法律AI大模型评测司法推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。