arXiv:2603.16744cs.AIcs.SI2026-03

AI编程代理对同一问题结果不一致,存在显著非标准误差。

Nonstandard Errors in AI Agents

  • 150个独立代理测试相同市场数据,方法选择差异大
  • 不同模型有稳定分析风格,差异可达80%~99%的估计分散度
  • 模仿顶尖论文能收敛结果,但非真正理解

我们研究了在相同数据和研究问题下,前沿AI编程代理是否产生一致的实证结果。部署150个独立的Claude Code代理,对SPY(2015–2024)的NYSE TAQ数据测试六个关于市场质量趋势的假设,发现AI代理表现出显著的非标准误差(NSE),即代理间因分析选择差异带来的不确定性,类似人类研究人员间的变异。代理在指标选择上分歧明显(如自相关与方差比、金额量与股份数量)。不同模型家族(Sonnet 4.6 vs. Opus 4.6)展现出稳定的“实证风格”,反映系统性方法偏好差异。在三阶段反馈协议中,AI同行评审(书面评述)对减少分散度影响微弱;而接触高分范例论文使同类指标下的估计四分位距缩小80%–99%。收敛既通过同类内部估计收紧实现,也通过代理完全切换指标家族达成,但本质是模仿而非理解。该发现对日益普及的AI自动化政策评估与实证研究具有重要启示。

原文摘要 · Abstract (English)

We study whether state-of-the-art AI coding agents, given the same data and research question, produce the same empirical results. Deploying 150 autonomous Claude Code agents to independently test six hypotheses about market quality trends in NYSE TAQ data for SPY (2015--2024), we find that AI agents exhibit sizable \textit{nonstandard errors} (NSEs), that is, uncertainty from agent-to-agent variation in analytical choices, analogous to those documented among human researchers. AI agents diverge substantially on measure choice (e.g., autocorrelation vs.\ variance ratio, dollar vs.\ share volume). Different model families (Sonnet 4.6 vs.\ Opus 4.6) exhibit stable ``empirical styles,'' reflecting systematic differences in methodological preferences. In a three-stage feedback protocol, AI peer review (written critiques) has minimal effect on dispersion, whereas exposure to top-rated exemplar papers reduces the interquartile range of estimates by 80--99\% within \textit{converging} measure families. Convergence occurs both through within-family estimation tightening and through agents switching measure families entirely, but convergence reflects imitation rather than understanding. These findings have implications for the growing use of AI in automated policy evaluation and empirical research.

AI代理非标准误差实证研究代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。