arXiv:2605.21076cs.CL2026-05

用大模型自动评德国法律考试,提升评分效率与可扩展性。

GradeLegal: Automated Grading for German Legal Cases

  • 通过提示工程逐步加入样例和评分标准,提升模型评分准确性。
  • 在公法领域模型评分与专家一致度达0.91(QWK),刑法仅0.60。
  • 集成多个模型效果更优,适合对评分可靠性要求高的场景。

德国法律考试评分面临考生数量增长与合格阅卷人短缺的矛盾,导致反馈延迟并形成瓶颈。由于国家考试成绩直接影响职业发展,该任务具有高重要性。然而,现有研究缺乏对法律考试自动化评分的有效方法系统性探索。本文评估27个私有及开源大语言模型在刑事与公法领域的自动评分能力,对比逐步添加样例、评分细则等任务信息的提示策略。结果表明,采用推理导向的模型,在提供样例和评分标准时,公法评分与专家一致度可达0.91(QWK),而刑事法仅为0.60,说明刑事法评分更具挑战性。此外,模型集成可使一致性提升0.15,优于表现最好的单一模型。研究强调有效提示设计与模型选择对实现可靠法律考试评分至关重要。

原文摘要 · Abstract (English)

Grading German legal exam solutions faces growing volumes and a shortage of qualified graders, delaying feedback and creating a bottleneck. At the same time, it is a high-stakes expert task, since state exam grades strongly influence career outcomes in Germany. Despite this practical relevance, literature lacks systematic studies on effective methods for grading legal exams. To address this gap, we investigate whether large language models (LLMs) can support the automated grading of German legal case solutions in criminal and public law, thereby enabling scalable feedback and student self-testing. We present a systematic evaluation of 27 proprietary and open-source LLMs, benchmarking prompting strategies that incrementally add task-related information, such as a sample solution and a grading rubric. Using quadratic weighted kappa (QWK), reasoning-oriented LLMs can approximate expert grading in public law when given a sample solution and a grading rubric (up to 0.91), compared to 0.60 in criminal law, suggesting a harder grading task in criminal law. Beyond single-model grading, ensembling improves agreement by up to 0.15 over its best member and can offer an alternative to stronger closed-source single models. In addition, our findings suggest that effective prompt design and model selection are necessary for reliable LLM-based grading of legal exams.

法律AI大模型自动评分德语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。