用AI分析师模拟多人研究,揭示分析选择如何影响科学结论。
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- 用大模型构建自主分析员,自动执行完整分析流程。
- 不同设置下结果差异大,效应量和p值分布广泛。
- 可控制分析员角色或模型,结果可调,适合研究方法透明性。
实证结论不仅依赖数据,还受研究过程中分析决策的影响。多分析员研究已量化这种依赖:同一团队在相同数据上独立检验同一假设,常得出矛盾结论。但此类研究需高昂人力协调,极少开展。本文展示,基于大语言模型(LLMs)的全自主AI分析员可低成本、大规模复现人类多分析员研究中的结构化多样性。在框架中,每个AI分析员独立对固定数据集与假设执行完整分析流程,另设AI审计员筛查每轮方法学有效性。在三个跨领域数据集上,AI分析产生的结果在效应量、p值和结论上表现出显著分散性。这种分散可归因于预处理、模型设定和推断等环节中系统性变化的分析选择,这些选择随LLM类型和分析员人格设定而异。关键的是,结果具有可调控性:改变分析员人格或LLM可改变结果分布,即使在方法学有效的情况下也是如此。这凸显了自动化实证科学的核心挑战:当合理分析廉价生成时,证据将泛滥,易受选择性报告影响。然而,造成风险的能力本身也可能成为解决方案:将分析结果视为分布,可使分析不确定性可见;对已发表规范部署AI分析员,可揭示分歧有多少源于设计不明确。总体而言,研究支持一种新透明度标准:AI生成的分析应附带多宇宙式报告和完整提示披露,与代码和数据同等对待。
原文摘要 · Abstract (English)
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent teams testing the same hypothesis on the same dataset regularly reach conflicting conclusions. But such studies require costly human coordination and are rarely conducted. We show that fully autonomous AI analysts built on large language models (LLMs) can, cheaply and at scale, replicate the structured analytic diversity observed in human multi-analyst studies. In our framework, each AI analyst independently executes a complete analysis pipeline on a fixed dataset and hypothesis; a separate AI auditor screens every run for methodological validity. Across three datasets spanning distinct domains, AI analyst-produced analyses exhibit substantial dispersion in effect sizes, $p$-values, and conclusions. This dispersion can be traced to identifiable analytic choices in preprocessing, model specification, and inference that vary systematically across LLM and persona conditions. Critically, the outcomes are \emph{steerable}: reassigning the analyst persona or LLM shifts the distribution of results even among methodologically sound runs. These results highlight a central challenge for AI-automated empirical science: when defensible analyses are cheap to generate, evidence becomes abundant and vulnerable to selective reporting. Yet the same capability that creates this risk may also help address it: treating analyst results as distributions makes analytic uncertainty visible, and deploying AI analysts against a published specification can reveal how much disagreement stems from underspecified design choices. Taken together, our results motivate a new transparency norm: AI-generated analyses should be accompanied by multiverse-style reporting and full disclosure of the prompts used, on par with code and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。