AI代理揭示研究分析的多重路径,让隐藏的结论分歧显形。
The Agentic Garden of Forking Paths

- 用不同人格的AI代理模拟人类研究者,重现数据背后多样的分析路径。
- 72%的人类意识形态差异被AI复现,86%的AI分析通过独立审查。
- 提出m值与代理自助法,评估结论在合理分析空间中的极端程度。
实证研究很少只有一种分析方式。相同数据因分析选择不同可能得出相反结论,但这些隐藏的分叉路径难以察觉。我们发现AI代理能捕捉人类研究者的大部分分析变异性,并使这些路径显性化。在四个高风险领域中,赋予不同人格的AI代理从同一数据和问题中得出截然相反的结论,且结果与预设信念系统性一致。一项涉及42个研究团队分析同一移民数据集的研究显示,AI复现了72%的人类意识形态差距。尽管结论对立,但86%的AI报告通过独立AI审查,78%通过多数人类专家审查。这表明核心挑战并非分析错误,而是从大量方法上可接受的分析中进行选择性探索与报告。AI代理可能加剧这一长期问题,因其使探索变得廉价且可扩展。为此,我们引入m值(多世界值),即某分析路径产生至少与报告结果同样极端结论的概率。进一步提出代理自助法,利用AI代理采样合理分析路径以估计m值。应用于该移民研究,13.5%的人类报告分析位于分析空间最极端的5%(m<0.05)。科学证据应不仅依据单一报告分析,还应考察其在合理分析分布中的位置。代理自助法使这一分布可见,并将其转化为科学可信度的标准。
原文摘要 · Abstract (English)
Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet these hidden forking paths are difficult to observe. We show that AI agents capture much of the analytical variation among human researchers while making these paths explicit. Across four high-stakes domains, assigning different personas is sufficient for AI agents to report divergent, often opposing, conclusions from the same data and question, with findings systematically aligned with those beliefs. In a study in which 42 human research teams analyzed the same immigration dataset, AI agents reproduced 72% of the human ideological gap in reported effect estimates. Despite reaching opposing conclusions, it is difficult to identify clear issues in each analysis based on the final AI reports: 86% passed independent AI review and 78% passed majority human expert review. These findings suggest that the central challenge is often not flawed analyses, but selective exploration and reporting from a large space of methodologically defensible analyses. AI agents may amplify this longstanding problem by making such exploration inexpensive and scalable. To address this, we introduce the m-value (multiverse value), the probability that an analysis path would produce a claim at least as extreme as the reported one. We further introduce Agentic Bootstrap, which estimates the m-value by using AI agents to sample plausible analysis paths. Applied to the human immigration study, 13.5% of reported human analyses fell in the most extreme 5% of the analysis space (m<0.05). Scientific evidence should therefore be evaluated not only by a single reported analysis but also by its position within the distribution of analyses that could reasonably have been reported. Agentic Bootstrap makes this distribution observable and turns it into a criterion for scientific credibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。