让AI仅凭论文方法描述和原始数据复现社科研究结果,验证可重复性。
Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results

- 从论文中提取结构化方法,隔离代码与结果进行自主复现
- 在48篇论文上实现大部分结果复现,但模型表现差异大
- 能定位失败原因,既找AI错误也暴露论文描述不足
现有工作利用大模型代理在有数据和代码的情况下复现社会科学研究结果。本文拓展这一范围:能否仅凭论文的方法描述和原始数据完成复现?我们构建了一个代理复现系统,从论文中提取结构化方法,严格隔离信息——代理不接触原代码、结果或论文全文,实现逐单元的确定性输出比对,并通过误差归因分析追踪偏差源头。在48篇经人工验证可复现的论文上,测试四种代理架构与四种LLM,发现代理可基本恢复已发表结果,但模型、架构及论文间表现差异显著。根因分析表明,失败既源于代理自身错误,也源于论文方法描述不充分。
原文摘要 · Abstract (English)
Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results given only a paper's methods description and original data? We develop an agentic reproduction system that extracts structured methods descriptions from papers, runs reimplementations under strict information isolation -- agents never see the original code, results, or paper -- and enables deterministic, cell-level comparison of reproduced outputs to the original results. An error attribution step traces discrepancies through the system chain to identify root causes. Evaluating four agent scaffolds and four LLMs on 48 papers with human-verified reproducibility, we find that agents can largely recover published results, but performance varies substantially between models, scaffolds, and papers. Root cause analysis reveals that failures stem both from agent errors and from underspecification in the papers themselves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。