大模型评分常不稳定,同一问题多次评估结果可能翻转,需多轮投票才可信。
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation

- 用50次重复实验发现,大模型对相同问题的判断有13.6%会翻转
- 多数情况下评分差距小,但依然频繁选出胜者,缺乏实质依据
- 建议采用多轮评分、随机排序和公开不确定性,避免单次打分误导
LLM-as-a-Judge现被广泛用于模型输出排序、训练奖励模型及构建公开排行榜,但其运行间可靠性仍不明确。本研究在29个任务(涵盖10类)上,使用两个OpenAI裁判模型(GPT-4o-mini 和 GPT-4.1-mini)进行50次成对比较与50次单点评分,辅以温度与提示敏感性分析。结果显示,成对偏好平均13.6%翻转,28%的问题超过20%翻转率,最高达56%。GPT-4o-mini还存在显著首位置偏差(72%选择第一项,p=0.024)。点均分差仅为0.19–0.36(10分制),总体不显著,说明评委常在无明显差异时仍选胜者。跨模型一致性仅76%(κ=0.51),语义等效提示模板导致25%案例改变多数意见,确定性解码虽降低波动但无法消除不一致。可靠性曲线显示,平均需11次重复才能以95%概率恢复50次基准结论,高方差问题需15次。这表明单次评估噪声过大,应常规采用多轮聚合、位置随机化与显式不确定性报告。因两模型同属一厂商,跨供应商复现仍是关键下一步。
原文摘要 · Abstract (English)
LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials and 50 pointwise trials per question, supplemented by temperature and prompt-sensitivity ablations. Across judges, pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one question reaching 56%. GPT-4o-mini also exhibits a significant first-position bias (72% A-majority, p = 0.024). At the same time, mean pointwise score gaps are small (0.19--0.36 on a 10-point scale) and not statistically significant in aggregate, producing a pairwise--pointwise gap: judges frequently choose a winner even when their own scalar scores provide little evidence of a meaningful quality difference. Beyond within-judge instability, cross-judge agreement is only 76% ($κ= 0.51$), semantically equivalent prompt templates change majority outcomes in 25% of tested cases, and deterministic decoding reduces but does not eliminate inconsistency. A reliability curve analysis shows that, in our dataset, 11 repeated trials are needed for a majority vote to recover the 50-trial reference verdict with 95% probability on average, rising to 15 for high-variance questions. These findings suggest that single-trial LLM judging is often too noisy for high-stakes evaluation, and that multi-trial aggregation, position randomization, and explicit uncertainty reporting should be standard practice. Because both judges are from a single provider, cross-provider replication remains an important next step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。