arXiv:2605.28183cs.CLcs.AI2026-05中稿 · EMNLP

评测大模型在德国法律中的归入式推理能力,构建了首个专项基准数据集。

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

论文配图:BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
图 1 · 摘自论文原文
  • 构建包含596个案例题和531个教义推理题的德国法律基准数据集。
  • 闭源大模型在三类任务中均领先,人机协作解题优于纯人工。
  • 验证了大模型评卷与人工评分高度一致,适合法律AI评估场景。

我们提出BenGER(德国法律基准),用于评估大模型在德国法律中的归入式推理能力。该数据集包含596个模拟考试风格的自由文本法律案例题,覆盖多个法学教育层级,以及531个简短教义推理题。数据集还包含受控验证子集,包含在无辅助和人机协同创作条件下由人类限时撰写的真实答案。我们对12种主流大模型系统进行了评估,采用与评分标准对齐的LLM作为裁判,并通过多评审者人类评分层(每份答案三次盲审,六组裁判与人类群体对比)进行交叉验证。闭源旗舰模型在所有三个语料库中均位居榜首;人机协同创作显著优于纯人工;大模型裁判与人类评分的相关性达皮尔逊r=0.76,克朗巴赫k=0.60。系统排名在不同裁判组间稳定,且来自独立提供方的两名裁判通过了Calderon单评审替代标准。

原文摘要 · Abstract (English)

We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The dataset combines 596 exam-style free-text legal case tasks across multiple levels of legal education and 531 short doctrinal reasoning tasks. It includes a controlled validation subset of timed human-written solutions under both unaided and human-AI co-creation conditions. We evaluate 12 contemporary LLM systems - closed flagship, efficiency-oriented, and open-weight - with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer (three blind reviews per solution, six judge families benchmarked against the human pool). Closed-flagship systems lead the leaderboard across all three corpora, human-AI co-creation measurably improves on unaided human work, and the LLM judge tracks human grading at Pearson r=0.76 and Cohen's k=0.60. System rankings are stable across judge families and two judges from independent providers clear the Calderon single-reviewer replacement bar on human-authored solutions.

法律AI大模型评测德国法归入推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。