arXiv:2603.11001cs.CYcs.AI2026-03

用随机实验评估AI对人类能力提升效果时,面临系统变化快、用户水平不一等挑战。

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

  • 通过专家访谈梳理出评估AI影响的三大方法难题
  • 发现快速迭代的AI系统会破坏实验有效性,影响结论可信度
  • 适合政策制定者和研究者参考,提升治理决策的科学性

人类提升研究(human uplift studies)通过随机对照试验(RCT)等方法评估AI对人类表现的影响,正日益为前沿AI治理与部署决策提供依据。尽管此类方法在其他领域稳健,但其在前沿AI系统中的适用性尚未充分探讨,尤其是在高风险决策场景下。本文基于对16位在生物安全、网络安全、教育和劳动力等领域有经验的专家的访谈,揭示了标准因果推断假设与前沿AI系统特性之间的持续张力。快速演进的AI系统、不断变化的基准线、用户能力差异及现实环境的开放性,均对研究的内部、外部和建构效度构成挑战,影响结果解释与实际应用。本文贡献包括:(1)系统归纳人类提升研究中的方法论难题,并按其与大语言模型(LLM)系统的相关性进行分类;(2)将挑战对应到可操作的解决方案。通过整合专家意见,旨在明确人类提升证据的解释边界,确保评估实践与所支持决策相匹配,并推动人工智能治理的协同方法基础。

原文摘要 · Abstract (English)

Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions. While RCT methods are robust in other fields, their interaction with the distinctive properties of frontier AI systems remains underexamined, particularly when results are used to inform high-stakes decisions. We present findings from interviews with 16 expert practitioners with experience conducting human uplift studies in domains including biosecurity, cybersecurity, education, and labor. Across interviews, experts described a recurring tension between the standard causal inference assumptions upon which human uplift studies rely and the object of study itself. Rapidly evolving AI systems, shifting baselines, heterogeneous and changing user proficiency, and porous real-world settings strain assumptions underlying internal, external, and construct validity, complicating the interpretation and appropriate use of uplift evidence. We contribute (1) a synthesis of methodological challenges in human uplift studies, mapped to risks to study validity and classified by their degree of specificity to large language model (LLM) systems, and (2) a mapping from challenges to proposed solutions. By collating expert-identified challenges and solutions, we seek to clarify the interpretive limits and appropriate uses of human uplift evidence, to align evaluation practice with the decisions it informs, and to support more coordinated methodological foundations for AI governance.

AI治理随机试验评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。