arXiv:2605.02050cs.CYcs.AI2026-05

为人工智能评估的随机对照试验制定标准化框架。

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

  • 基于多学科经验提炼五大原则,构建可操作指南
  • 提出33条带理由和实施建议的评估准则
  • 帮助研究者设计实验、评审论文并推动行业标准

本文建立了一个用于标准化人工智能评估随机对照试验(RCT)的框架(有时称为人类提升研究)。借鉴软件工程、经济学、临床与健康科学、心理学等已有RCT传统学科的成熟实践,综合既有的效度框架和开放科学标准中的透明性、可重复性与可验证性,提炼出五个核心原则,并据此发展出33条针对AI评估场景的可操作指南。每条指南均包含要求、理由、实施说明和证据依据。该框架在三个方面发挥作用:作为研究设计工具、现有工作的评估量表,以及未来标准制定的蓝图。当前人工智能评估研究缺乏统一标准与共享术语,难以生成可累积、可比较、可用于政策决策的证据。本框架为此奠定基础,提供评价标准与共用概念语言,辅以具体行动指南。

原文摘要 · Abstract (English)

This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we synthesize five principles drawn from established validity frameworks and open-science standards on transparency, repeatability, and verification, which together serve as the conceptual foundation for 33 actionable guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases. We position the principles and guidelines as serving three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms. AI evaluation research currently lacks common standards and shared vocabulary for producing cumulative, comparable, policy-ready evidence. This framework is a contribution toward that foundation, providing evaluative criteria and a shared conceptual language alongside actionable guidelines.

评估标准随机对照试验AI伦理可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。