揭示大模型研究中低成本方法的隐藏缺陷,提供评估有效性的系统框架。
Validity Threats for Foundation Model Research

- 将大模型研究视为因果推断问题,用四类有效性分析方法
- 发现代理实验牺牲外部与建构有效性以换取内部统计可信度
- 指出单次训练设计受干预单元干扰,易被忽略的潜在偏差
控制实验是机器学习研究的核心,但在现代大模型规模下成本过高。社区转向低成本近似方法:代理实验、缩放定律、基于公开模型的观察研究,以及利用单次训练过程内变异的单次运行设计。本文认为,在计算预算限制下,任何节省都伴随有效性威胁——隐含且有时无法验证的假设,一旦失效,研究结论即可能无效。为此,我们提出一个评估框架,将大模型研究建模为因果推断问题,借鉴社会科学中的四种有效性标准(统计、内部、外部、建构)来评估不同研究策略。结果显示:代理实验以牺牲外部和建构有效性换取统计与内部有效性;观察研究面临混淆因素与效应异质性;单次运行设计则受处理单元间干扰困扰。该分析揭示了文献中未充分关注的若干有效性威胁。整体上,本框架为研究人员审视大模型研究设计中的有效性风险提供了实用工具。
原文摘要 · Abstract (English)
Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively expensive. Instead, the community increasingly relies on research strategies that approximate the ideal experiment at a fraction of the cost: proxy experiments and scaling laws, observational studies with publicly available models, and single-run designs that leverage variation within individual training runs. In this work, we argue that there is no free lunch when approximating large-scale experiments on a compute budget. Specifically, savings in compute come at the cost of validity threats -- hidden and sometimes untestable assumptions that, when violated, can invalidate research claims. To help navigate such threats, we propose an evaluation framework that casts foundation model research as a causal inference problem. Within this framework, we evaluate different research strategies through four types of validity adapted from the empirical social sciences -- statistical, internal, external, and construct validity. We find that each strategy comes with a characteristic validity profile: proxy experiments trade external and construct validity for statistical and internal validity; observational studies face confounding and effect heterogeneity; and single-run designs are strained by interference between treated units. This analysis reveals several validity threats that have received insufficient attention in the literature. Overall, our evaluation framework provides researchers with a practical toolkit for scrutinizing validity threats in foundation model research~designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。