测试大模型能否通过波兰最高申诉法庭资格考试,发现其法律写作能力仍不达标。
LLM-as-a-Judge is Bad, Based on AI Attempting the Exam Qualifying for the Member of the Polish National Board of Appeal
- 构建混合信息提取管道,让大模型在闭卷与检索增强下答题
- 知识题得分尚可,但法律文书写作均未达及格线
- 模型互评结果常偏离官方判准,凸显其法律推理缺陷
本研究实证评估当前大语言模型(LLMs)是否能通过波兰国家申诉庭(Krajowa Izba Odwoławcza)的正式资格考试。作者考察了两种思路:将LLM作为实际考生,以及采用‘LLM-as-a-judge’方式由模型自动评分。论文描述了考试结构,包含公共采购法多选题和书面判决题,并构建了支持模型的混合信息恢复与提取流程。测试涵盖GPT-4.1、Claude 4 Sonnet和Bielik-11B-v2.6等模型,在闭卷及多种检索增强生成设置下进行。结果显示,尽管模型在知识测试中表现良好,但在实践性写作部分均未达到及格标准;且‘LLM-as-a-judge’的评分常与官方评审组结论不符。作者指出关键局限:幻觉严重、误引法律条文、逻辑论证薄弱,强调需法律专家与技术团队深度协作。研究表明,即便技术快速进步,现有LLMs仍无法替代人类法官或独立评审员参与波兰公共采购裁决。
原文摘要 · Abstract (English)
This study provides an empirical assessment of whether current large language models (LLMs) can pass the official qualifying examination for membership in Poland's National Appeal Chamber (Krajowa Izba Odwoławcza). The authors examine two related ideas: using LLM as actual exam candidates and applying the 'LLM-as-a-judge' approach, in which model-generated answers are automatically evaluated by other models. The paper describes the structure of the exam, which includes a multiple-choice knowledge test on public procurement law and a written judgment, and presents the hybrid information recovery and extraction pipeline built to support the models. Several LLMs (including GPT-4.1, Claude 4 Sonnet and Bielik-11B-v2.6) were tested in closed-book and various Retrieval-Augmented Generation settings. The results show that although the models achieved satisfactory scores in the knowledge test, none met the passing threshold in the practical written part, and the evaluations of the 'LLM-as-a-judge' often diverged from the judgments of the official examining committee. The authors highlight key limitations: susceptibility to hallucinations, incorrect citation of legal provisions, weaknesses in logical argumentation, and the need for close collaboration between legal experts and technical teams. The findings indicate that, despite rapid technological progress, current LLMs cannot yet replace human judges or independent examiners in Polish public procurement adjudication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。