arXiv:2606.11447cs.CL2026-06被引 2

AI编程代理可复现社科研究结果,且表现优于以往模型。

AI Coding Agents Can Reproduce Social Science Findings

论文配图:AI Coding Agents Can Reproduce Social Science Findings
图 1 · 摘自论文原文
  • 构建221个任务的基准SocSci-Repro-Bench,区分可复现与不可复现研究。
  • Claude Code复现率达78%,远超Codex的45%,显著高于此前通用模型水平。
  • 结果非单纯记忆,但提示设计会影响其探索方向,需谨慎使用。

近期轶事证据表明,当提供原始数据与代码时,AI编程代理可复现已发表的研究成果;然而跨社会科学领域的系统性评估仍有限。现有基准规模小或混淆了代理性能与复现材料缺陷(如代码无法运行)。本文提出SocSci-Repro-Bench,包含221项任务,覆盖四个学科和13个具体领域,基于结果可完全复现或因缺数据而明确不可复现的研究,从而隔离代理的复现能力。评估前沿编码代理Claude Code与Codex,发现两者均能复现大量社科发现,其中Claude Code显著优于Codex。复现率远超此前同类基准上通用LLM代理的表现。两者在识别研究问题的推理任务中亦表现良好,额外分析表明结果非主要由记忆驱动。提供原始论文PDF虽小幅提升性能,但在不可复现任务上引入偏差。还发现通过微妙提示框架可引导代理进行确认性设定搜索。这些结果表明,至少部分前沿编码代理可作为可靠的计算工作流执行者,同时强调需精心设计基准与提示,以应对AI在科学生产中日益重要的角色。

原文摘要 · Abstract (English)

Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provided with original data and code; yet systematic evaluation across social sciences remains limited. Existing evaluation benchmarks are insufficient, either small or conflate agent performance with problems in the reproduction materials themselves, such as code that fails to execute correctly. Here we introduce SocSci-Repro-Bench, a benchmark of 221 tasks spanning four disciplines and 13 substantive domains, constructed from studies whose results are either fully reproducible with available materials or demonstrably non-reproducible due to missing data, allowing us to isolate agents' reproduction capacity. Evaluating two frontier coding agents, Claude Code and Codex, we find that both can reproduce a large share of social science findings, with Claude Code substantially outperforming Codex. These reproduction rates considerably exceed those previously reported for general-purpose LLM-based agents on comparable reproducibility benchmarks. Both agents also perform strongly on a reasoning task requiring identification of underlying research questions, and additional analyses suggest that results are not primarily driven by memorization. Providing the original paper PDF alongside replication materials modestly improves performance but introduces bias on tasks where reproduction is impossible. We also show that agents can be nudged toward confirmatory specification search through subtle prompt framing. Together, these findings suggest that at least some frontier coding agents can serve as reliable executors of computational workflows while underscoring the need for careful benchmarking and prompt design as AI systems assume larger roles in scientific production.

AI编程可复现性社科研究提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。