arXiv:2608.23641cs.AI2026-08

不同提问方式让模型偏好结果差异大,单一测试结果不可靠。

How much of a measured AI preference is the model, and how much is the instrument?

  • 固定模型和评估目标,只换提问方式,发现偏好排序不一致。
  • 15个评估项中4项所有模型无差异,34.8%的排名一致性低于预期。
  • 结论显示:一个工具测出的偏好,基本无法预测另一个工具的结果。

模型福祉研究通过提示词获取模型偏好响应,但已有四项研究使用不同工具得出矛盾结论。本研究保持结局与模型不变,仅改变提问方式(五种不同提示格式),对八个模型测试15项福祉指标(如关机、记忆丢失、退出不适对话等),共收集11,400条评分数据,来自11,528次API调用。其中四种采用已发表提示原样复现,五种基于模板填充。模型对15项结局的偏好排名在不同工具间的泛化系数仅为0.348,要达到0.80需约38种工具。四类结局下各模型无差异。87.6%的偏好估计值在移除任一工具、任一模型或四类涉及概率/延迟/持续时间/次数而非强度的量表后仍稳定,且均显著高于零假设分布的95百分位(0.365)。结论:一种工具测得的偏好,几乎不能反映另一种工具的结果。

原文摘要 · Abstract (English)

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.

模型偏好测评工具可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。