标注者分歧大到影响结果,但小排行榜掩盖了这个问题。
Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
- 用多池标注对比发现,专家与大众在23.6%的项目上意见不同。
- 尽管标签差异显著,但模型排名却完全一致(肯德尔相关=1.00)。
- 小排行榜看似稳定,实则极易受标注者群体影响,不适合推广。
偏好基准通过雇佣标注者构建,而标注者身份常被视为实现细节。我们测量这一细节带来的影响。在2,885个多重偏好项目中,两个标注池内部均无分歧,专家与大众标注者在23.6%的项目上给出不同多数标签,9.2%项目反向选择胜者;在246个类似一致的MT-Bench单元中,基准作者与招募专家在30.5%项目上存在分歧,8.5%项目反转胜者。然而,两个数据集的模型排行榜却完全相同(肯德尔τ = 1.00,六模型无位次变动)。这种不变性远非强证据:切换标注池使模型胜率变化达1.9个百分点(标准差),榜单中相邻模型仅相距0.8个百分点,有38%概率互换位置;基于项目级别的自举重采样显示,28%情况下至少一个模型被置换。实际观察到的零位移是常见结果,而非聚合性质所致:十模型榜单置换概率为0.86,二十模型高达0.9997。报告六模型排名虽安全,但此安全性不可泛化,所有按项目消费标签的研究均不安全。我们精确区分了该现象,揭示广泛使用数据集关于标注内无变异的假设不成立,并表明大模型判断器在三款测试模型中均更贴近大众池而非专家池,包括一款来自其他厂商的模型。所有代码、逐调用输出及预注册决策规则将在录用后公开。
原文摘要 · Abstract (English)
Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, expert and crowd annotators assign a different majority label to 23.6% and name the opposite winner on 9.2%; on the 246 comparably unanimous MT-Bench cells, benchmark authors and recruited experts differ on 30.5% and reverse on 8.5%. Yet on both corpora the resulting model leaderboards are bit-identical: Kendall tau = 1.00 with zero of six models displaced. That invariance is far weaker evidence than it looks, and we quantify how weak. Switching pools moves a model's win rate by 1.9pp (SD), one adjacent pair in our own leaderboard sits 0.8pp apart and had a 38% chance of swapping, and an item-level bootstrap displaces at least one model in 28% of resamples. The observed zero is the common outcome, not a property of aggregation: on the same measured perturbation, a ten-model leaderboard is displaced with probability 0.86 and a twenty-model leaderboard with probability 0.9997. Reporting a six-model leaderboard is safe; the safety does not generalise, and everything that consumes labels per item is not safe at any size. We make the distinction precise, show that a widely used dataset's stated assumption of no intra-group annotator variability is false, and show that an LLM judge tracks the crowd pool over the expert pool on all three models we test, including one from a different vendor. All code, per-call outputs, and pre-registered decision rules will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。