语言模型之间比与人类更一致,且这种一致性不依赖于提示或厂商。
Language Models Agree With Each Other, Not With Readers
- 用真实读者标注数据对比模型,避免人为设计偏差。
- 模型间重合度达8.7,远超人类间的4.1,且多数模型超越人类表现。
- 一致性源于模型本质特性,非提示或厂商差异,适合关注模型行为的研究者。
现有研究衡量语言模型同质化时,常以人工判断为基准,但这些判断本身是模型提示的产物,存在设计偏差。本文采用2,523个读者在120篇网页文档上的独立标注集合,其标注叠加默认关闭,代表真实阅读意图。通过计算大小匹配句集间的重叠度,并减去基于深度和长度分布重采样后的期望重叠,建立零假设校准。所有随机基线对均落在零附近(±0.006)。中位数文档中,每名读者标记14句(共70句),两人共享4.1句;两模型共享8.7句。18种模型配置(11家厂商、3国、两种权重)中,153对模型的中位一致性为+0.093,高于人类基准的+0.040,99对完全高于人类区间。两个前沿模型达+0.203,是GPT-4o自评值的两倍。该现象非确定性、提示、流程、厂商或路由所致,且呈规模梯度:最小模型与人类一致。无模型显著优于读者,且在相同深度与长度下,表面特征无法区分选择模式。差异程度依赖流程,排序则非:模型被裁剪至最锐利集,读者则是随机抽样,即便统一模糊处理,差距仍存。四款后期发布模型经预先预测测试,均未达到人类区间。多模型模拟群体并非多个独立群体。
原文摘要 · Abstract (English)
Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。