用预注册机制防止大模型研究中的假阳性陷阱。
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
- 预注册分析方案和未来可用模型,只在首个符合条件的模型上执行。
- 在20个模型上测试,73.9%~72.7%的假阳性攻击被成功阻断。
- 适合注重可复现性的大模型实证研究者使用。
大语言模型(LLMs)被广泛用于生成、分类和标注数据,这些数据随后用于下游假设检验。然而,基于LLM的研究极易产生p-hacking:研究者可通过调整提示词、解码参数或输出格式,直到获得期望结果。本文提出一种缓解策略:预注册实验方案及一批候选未来模型,随后在首个符合条件的模型发布后立即运行确认性分析。由于该模型在承诺时并不存在,无法针对性地进行操控;且针对一个模型有效的配置通常不适用于下一个模型。我们在两个真实值已知的任务上评估该协议,在4家厂商共20个模型和11种分析配置下,该协议成功阻止了73.9%和72.7%的假阳性转移。额外压力测试显示其效果依然显著。最后,我们以身作则,按协议预注册实验,结果显示在7个曾成功攻击前一模型的配置中,有6个在新模型上失败。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p-hack: a researcher can tune the prompts, decoding parameters, or output format until a desired result is reached. We propose a protocol to mitigate p-hacking in LLM-based research: preregistering the experiment and eligible models, and then running it on the first eligible LLM that is released after the preregistration. The researcher finalizes the procedure on current models, preregisters the analysis plan together with a set of eligible future models, and runs the confirmatory analysis on the first eligible model released afterward. Because this model does not exist at commitment time, it cannot be hacked against; furthermore, configurations that hack one model frequently do not transfer to the next. We evaluate the protocol on two tasks whose true values are known. Across 20 models from four providers and 11 LLM-analysis configurations, the protocol would have blocked successful transfer of the p-hack in 73.9% and 72.7% of cases in the two tasks. Additional analyses reveal that mitigation remains substantial under several stress tests. Finally, putting money where our mouth is, we followed our own protocol and preregistered our experiment. The preregistered experiment confirmed the protocol's effectiveness: out of the 7 configurations that hacked the prior model, the hacking failed to carry over in 6 configurations on the first eligible model released afterward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。