研究发现大模型评卷与人类专家存在不对称分歧,尤其在规则模糊处。
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
- 用同一组试题测试26个大模型和3位人类考官,检验评分一致性
- 多数题目大模型与多数人类考官一致,但5个模糊题中仅1个模型接近少数人类答案
- 大模型群体呈现系统性趋同,难以复现人类中的少数观点,适合评估主流判分
自然语言处理中的自由形式法律作文评估常将专家评分一致性视为单一上限,并以大模型与该上限的吻合度作为其稳定性的证据。本文通过相同输入协议,在泰国律师资格考试中测试这两个假设:三位由巴协会培训的考官(A、B、C)与一个由26个大模型组成的评审团对15份交叉评分的作答(来自4个输入:题目、官方评分标准、标准答案、考生作答)进行评分。主要发现为非对称性:在10个评分标准明确要求双维度评分的题目中,29名评审者全部集中在紧密区间内,达成普遍一致;而在其余5个未规定正确答案若遗漏关键法条引用时如何评分的题目中,人类考官分裂为两个合理解读——B/C多数位于高分段(6-8分),而考官A则位于低分段(1-2分)。大模型群体并未对称分裂:22/26个模型落在B/C的争议区间,3个处于规则空白的中间地带,仅1个(GPT-5.4 Nano)接近但未持续进入A的区间。我们26个模型中无一复现人类少数观点。倾向B/C方向的集群涵盖所有模型规模、供应商及价格层级。通过三模型锚点子集(Claude 4.6 Opus、Gemini 3.1 Pro、GPT-5.4 Pro)开展确定性探测、输入消融与置信区间分析,锚点面板在15个题目上达到α=0.77,远高于人类面板α=0.36。高α值反映的是对多数意见的系统性趋同,而非对两种观点的平衡再现;若以最大化与人类参考组一致性的标准选择大模型裁判,将不可避免继承这种不对称性。
原文摘要 · Abstract (English)
Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling as evidence of judge stability. We test both assumptions on the Thai bar examination through an identical-inputs protocol: three Bar Council-trained examiners (A, B, C) and a 26-LLM judge panel score the same 15 cross-graded answers from the same four inputs (question, official Bar Council grading regulation, gold answer, candidate answer). The headline finding is asymmetric. On 10 of 15 cells where the rubric prescribes both axes, all 29 raters converge in a tight band: panel agreement is universal. On the remaining 5 cells where the rubric does not prescribe how to grade a correct final answer that omits a decisive statutory citation, the human panel splits between two coherent readings (B/C majority at the upper rubric band, score 6-8; A minority at the lower band, score 1-2). The LLM judge population does not split symmetrically: 22 of 26 LLMs score in or near B/C's contested band, 3 sit in the regulation-silent middle gap, and only 1 (GPT-5.4 Nano) approaches A's band without consistently scoring within it. Zero LLMs in our 26-judge panel reproduce the minority human reading on the contested cells. The B/C-direction cluster spans every model size, vendor, and price tier we tested. An instrumented three-LLM anchor sub-panel (Claude 4.6 Opus, Gemini 3.1 Pro, GPT-5.4 Pro) carries determinism probes, input ablations, and bootstrap CIs, and reaches anchor panel $α= 0.77$ on the 15 cells against human-panel $α= 0.36$. The high LLM-panel $α$ reflects systematic convergence on the majority reading rather than balanced reproduction of both readings; a benchmark that selects its LLM judge by maximising agreement with a human reference panel will inherit this asymmetry by construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。