arXiv:2608.27309cs.CLcs.AI2026-08

LLM评估中的双重差分法可能因评分上限导致虚假效应,需警惕。

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

  • 在受限评分尺度下,双重差分法受截断影响,无法准确识别真实偏好差异
  • 实验显示名义显著的交互效应中,85%可由评分下限与严重度偏移解释
  • 适用于审查大模型评估结果可信性的研究人员和算法审计者

对大型语言模型(LLM)评判者的审计通过对比匹配条件来验证偏差,最强设计采用双重差分:先在单个样本内比较两个候选回答,再跨被操纵属性进行第二次差分,并基于有界评分尺度读取结果。我们证明该终点在报告的尺度上并非可识别。双重差分的每一项均受自身占比的截断影响,因此观测统计量混淆了偏好差异与衰减差异:当两回答同时受到严重度偏移影响时,只要它们距边界距离不等,就必然产生交互效应,而优质刺激恰好位于此位置。我们在一项预注册的冻结教学判别器审计中展示了这一失效,该审计在首次990次调用前已封存。注册的主要终点——学习者身份对支架偏好影响——为零:+0.085分(95% BCa [-0.167, +0.353],p=0.684)。审计中唯一名义上显著的交互效应为+0.378(p=0.002),但并非源于偏好:一个不含差异偏好的构造仅凭观测到的严重度偏移与评分下限即可复现其中79%至85%。我们以闭式形式推导出该机制,并表明其贡献可从审计自身评分中测量。

原文摘要 · Abstract (English)

Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.

大模型评估双重差分评分偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。