arXiv:2608.07517cs.HCcs.CL2026-08

大模型看网页截图预测A/B测试结果,多数不准,但自信判断可识别可信结果。

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

  • 用大模型分析网页截图预测A/B测试胜者,但整体准确率仅0.14(低相关)
  • 模型自信时的判断在显著标签上达kappa=0.31,覆盖49%的可信测试
  • 专家与模型共识高但与真实结果无关,集体意见不等于证据

能否仅凭网页截图让多模态大模型预测真实A/B测试的胜者?我们基于六周预注册实验给出最完整答案:多数情况下不能——例外可提前识别。在330个真实转化测试中,Gemini 3 Flash裁判的Cohen's kappa为0.14;而在可信(统计显著)的半数标签上,证据仍不充分(kappa=0.11,置信区间包含零)。我们发现,领先客户体验机构数据库中44%的“真实标签”来自非显著测试,且模型更认同不可靠标签——这是标签者与模型共享的先验,而非预测能力。所有标准优化手段(2.8倍成本的前沿模型、提示重设计、刺激保真度、变化类型先验)均未通过预注册检验。但模型自信判断表现不同:投票阈值筛选出49%覆盖范围,其在显著标签上的kappa达0.31。我们直接测量机制:不同模型或提示的裁判彼此一致性高达kappa=0.74–0.88,但与真实结果一致率仅约0.2,16人评审团实际仅相当于2个有效独立判断。人类复现结果:15名CRO专家内部一致性kappa=0.53,但与真实结果一致率接近随机(kappa~0)。共识(无论人或模型)可重复、具说服力,但不是证据。我们公开所有预注册、锁定门禁、负面结果、统计工具、人类响应及带有证据层级的声明清单。

原文摘要 · Abstract (English)

Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the "ground-truth" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly -- judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only ~0.2, so a 16-vote panel carries about 2 effective independent votes -- and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa ~ 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.

A/B测试大模型评估可信预测判别校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。