arXiv:2606.15610cs.CLastro-ph.IM2026-06被引 1

给大模型当裁判的行为做体检,发现它们有隐藏偏见和测量误差。

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

论文配图:LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
图 1 · 摘自论文原文
  • 提出裁判数据表,量化模型在无输入时的虚假偏好
  • 实测发现不同模型存在位置偏差、响应混乱等测量缺陷
  • 适合评估大模型评判能力的研究者与开发者使用

LLM-as-a-judge系统常用于开放任务评估,替代成本高且难复现的人类偏好标注。但当前多以准确率、胜率或一致性报告其表现。我们主张应将裁判视为测量仪器。本文提出裁判数据表协议,测量真实真空下的暗电流、对同质表面变化的稳定性、位置性错误偏好、可控质量阶梯上的目标敏感度,以及平局指令引发的判别标准。方向-稳定性分解显示,看似稳定的Δ0偏好可能源于表面响应或隐藏的位置偏差。三模型案例研究中,Llama-3.1-8B表现出高暗电流和呈现冲突的Δ0行为;Qwen2.5-14B真空洁净但混合稳定与位置过敏;Qwen2.5-32B真空洁净、交叉敏感度低、位置误判少。严格平局规则消除Qwen32B的Δ0误判,但会将微弱Δ1信号归入平局,同时保留Δ5敏感性。结果表明,提示词改变的是判别标准而非分辨能力。本文不证明下游机制假设,贡献在于为评估前建立计量协议。

原文摘要 · Abstract (English)

LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce. Yet these judges are often reported as scalar accuracy, win-rate, or agreement devices. We argue that a judge should instead be reported as a measurement instrument. We introduce a Judge Datasheet protocol that measures dark current under true-vacuum inputs, stable cross-sensitivity to same-quality surface variation, positional false preference, target sensitivity on a controlled quality ladder, and the criterion or operating point induced by tie instructions. The direction-stability decomposition reveals that apparent Delta0 preference can be stable surface response or disguised position bias. In a three-judge open-weight case study, Llama-3.1-8B shows high dark current and presentation-conflicted Delta0 behavior, Qwen2.5-14B is vacuum-clean and target-sensitive but mixes stable and positional over-discrimination, and Qwen2.5-32B is vacuum-clean with low stable cross-sensitivity and low positional false preference. A strict tie criterion eliminates Qwen32B Delta0 false preference but absorbs marginal Delta1 target signals into ties while preserving Delta5 sensitivity. The results show that prompting moves the criterion, not the resolution. We do not claim that the downstream mechanism hypothesis that motivated this work is confirmed; the contribution is a metrological protocol for measuring the measuring device before downstream claims are made.

大模型评测裁判机制心理测量模型偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。