LLM当裁判时,换人就会改分数,可靠性存疑。
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

- 用不同规模和版本的LLM做裁判,相同回答得分会变
- 只有从Qwen3 1.7B升级到4B时得分提升稳定,其他升级无效
- 建议报告时附上偏差探测、误差相关性等审计信息
当候选回答固定时,仅因更换评估者(LLM-as-judge)会导致评分变化。我们将其视为测量有效性问题。在四个判断数据集上,比较了两种常见升级路径:将Qwen3稠密裁判从1.7B参数扩展至32B,以及使用MiniMax M2-M2.7发布的API。主要发现是:裁判升级不可互换——仅从Qwen3 1.7B升至4B带来稳健增益,而MiniMax相邻版本无明显改善。更强的裁判虽能缓解位置与冗长偏倚,但无法消除。重复采样组成的评审团在误差相关时贡献有限。结构化辩论可显著改变判决结果,但若无解析器与回退日志,无法确认改变源自讨论。我们主张LLM-as-judge报告应包含数据集切片、偏差探测、误差依赖估计及协议审计轨迹。
原文摘要 · Abstract (English)
An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment datasets, we compare two upgrade paths available in practice: scaling Qwen3 dense judges from 1.7B to 32B parameters and moving across MiniMax M2-M2.7 released APIs. The main pattern is that judge upgrades are not interchangeable: only Qwen3 1.7B to 4B gives a robust adjacent gain, while MiniMax adjacent releases do not. Stronger judges reduce but do not remove position and verbosity bias. Repeated-sample juries add little when errors are correlated. Structured debate can move decisions substantially, but without parser and fallback logs those shifts cannot be attributed to deliberation. We argue that LLM-as-judge reports should include dataset slices, bias probes, error-dependence estimates, and protocol audit trails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。