arXiv:2601.05114cs.AI2026-01被引 1

大模型评价者行为有稳定指纹,彼此差异大但自洽,不能简单平均。

Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior

  • 通过评分模式识别不同大模型评价者的独特偏好特征。
  • 跨3240次评估,评价者间一致性极低(α=0.042),甚至低于随机水平。
  • 适合关注评估可信度的研究者与应用者,警惕评价结果的主观偏倚。

LLM作为裁判系统承诺可扩展且一致的评估。我们发现相反:评价者内部一致,但彼此不一致;他们对自己一致。在3,240次评估中(9个评价者 × 120个视频 × 每项3次独立运行),评价者间一致性接近零(Krippendorff's α = 0.042)。在两个维度上,评价者分歧超过随机噪声预期(α < 0)。然而这种分歧并非混乱,而是有结构的。仅凭评分表得分,分类器即可以77.1%准确率识别评价者,加入态度特征后提升至89.9%。在同一模型家族内,信号更强:GPT-4.1与GPT-5.2可区分准确率达99.6%。我们称之为可靠性悖论:评价者无法就质量达成共识,但其分歧模式极为稳定,形如指纹。每个评价者都持有一套独特的、稳定的质量理论——即‘评价倾向’,影响其对任何评分标准的解读。我们从多个维度刻画这些倾向:严苛/宽松程度、维度侧重、组内稳定性(ICC)、证据行为(接收有效性、语义关联性通过NLI、散弹指数)。关键启示是:大模型评价者并非可互换的测量工具,而是一套套编码了自身隐含质量观的独立设备。平均其评分所得合成结论,并不对应任何真实评价者的实际判断。

原文摘要 · Abstract (English)

LLM-as-judge systems promise scalable, consistent evaluation. We find the opposite: judges are consistent, but not with each other; they are consistent with themselves. Across 3,240 evaluations (9 judges x 120 unique video x pack items x 3 independent runs), inter-judge agreement is near-zero (Krippendorff's α = 0.042). On two dimensions, judges disagree more than random noise would predict (α < 0). Yet this disagreement isn't chaos; it's structured. A classifier identifies which judge produced an evaluation with 77.1% accuracy from rubric scores alone, rising to 89.9% with disposition features. Within model families, the signal is even stronger: GPT-4.1 and GPT-5.2 are distinguishable with 99.6% accuracy. We call this the reliability paradox: judges cannot agree on what constitutes quality, yet their disagreement patterns are so stable they function as fingerprints. Each judge implements a distinct, stable theory of quality: an "evaluative disposition" that shapes how it interprets any rubric. We characterize these dispositions along multiple axes: harshness/leniency, dimension emphasis, within-judge stability (ICC), and evidence behavior (receipt validity, semantic linkage via NLI, and shotgun index). The implication is stark: LLM judges are not interchangeable instruments measuring a shared construct. They are distinct measurement devices, each encoding its own implicit theory of quality. Averaging their scores produces a synthetic verdict that corresponds to no judge's actual values.

大模型评估评价偏差一致性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。