黑盒大模型评估系统稳定性存疑,相同请求多次运行结果不一致。
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
- 在预注册实验中验证评估工具可靠性,发现重复请求结果波动大。
- 同一请求重跑,相关性仅0.40(需0.90),次日重试也仅0.78(需0.99)。
- 适合关注大模型评估可信度的研究者与开发者参考。
语言模型评判器现用于筛选训练数据、评分生成内容并主导排行榜。其作为测量工具,依赖一个未明言的假设:相同请求发送至相同模型名,结果应稳定不变。本文通过两次预注册审计,在所有阈值预先固定的前提下,均未能验证该仪器的可靠性。在52,988次测试请求中,同一窗口内重复排名的斯皮尔曼相关性仅为0.400(要求0.90),字节完全相同的次日重试相关性为0.78(要求0.99),且执行记录已达上限。三个机制解释差异:标签到语义映射偏差强于信号本身;候选差距低于仪器噪声下限七个数量级;字节完全相同的输入返回不同排名,且精确排列读取会放大这种噪声。任何指标替换或采样策略均未修复问题。后续预注册实验表明:等待无改善(0.805对0.800,五天重复);切换提供商无效(四家共享低水平,中位数0.74–0.88,均无法由暴露元数据预测);自托管在批次无关内核上仅在服务器空闲时有效;在构造错误中,读出分离度反映错误类型而非大小。研究提炼出三层快照-身份阶梯、八项设计规则及报告检查清单;在约2%的调用量下进行试点即可提前暴露不可达门槛。所有结果基于共享服务基础设施上的外部行为测量。在共享端点上,模型名称不是静态仪器;预注册评估必须在冻结任何门禁前先测量其仪器可靠性。
原文摘要 · Abstract (English)
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。