arXiv:2507.18901cs.CL2025-07ACL被引 23

用AI自动评估社科论文可复现性,挑战大但有突破。

REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?

  • 设计真实场景的可复现性评估任务,覆盖多种数据格式与编程语言。
  • 现有AI agent最高准确率仅21.4%,表明自动化评估仍具挑战。
  • 提出REPRO-Agent模型,性能比最优现有模型提升71%,适合研究者参考。

评估社会科学论文的可复现性对提升研究严谨性至关重要,但人工评估成本高。随着智能体类AI系统的发展,我们尝试评估其自动化该任务的能力。然而,现有复现基准存在三方面不足:(1) 仅关注使用给定代码和数据复现结果,未评估其与论文的一致性;(2) 过度简化真实场景;(3) 缺乏数据格式与编程语言的多样性。为此,我们构建了REPRO-Bench,包含112个任务实例,每个对应一篇公开可复现报告的社会科学论文。智能体需基于原始论文PDF和复现包评估可复现性。该基准涵盖复杂度接近真实世界的端到端评估任务。我们在三个代表性AI智能体上测试,最佳表现者准确率为21.4%。基于实证分析,我们提出REPRO-Agent,将现有最佳模型准确率提升71%。结论指出需发展更先进的智能体以实现真实世界可复现性评估。REPRO-Bench已开源,地址:https://github.com/uiuc-kang-lab/REPRO-Bench。

原文摘要 · Abstract (English)

Assessing the reproducibility of social science papers is essential for promoting rigor in research processes, but manual assessment is costly. With recent advances in agentic AI systems (i.e., AI agents), we seek to evaluate their capability to automate this process. However, existing benchmarks for reproducing research papers (1) focus solely on reproducing results using provided code and data without assessing their consistency with the paper, (2) oversimplify real-world scenarios, and (3) lack necessary diversity in data formats and programming languages. To address these issues, we introduce REPRO-Bench, a collection of 112 task instances, each representing a social science paper with a publicly available reproduction report. The agents are tasked with assessing the reproducibility of the paper based on the original paper PDF and the corresponding reproduction package. REPRO-Bench features end-to-end evaluation tasks on the reproducibility of social science papers with complexity comparable to real-world assessments. We evaluate three representative AI agents on REPRO-Bench, with the best-performing agent achieving an accuracy of only 21.4%. Building on our empirical analysis, we develop REPRO-Agent, which improves the highest accuracy achieved by existing agents by 71%. We conclude that more advanced AI agents should be developed to automate real-world reproducibility assessment. REPRO-Bench is publicly available at https://github.com/uiuc-kang-lab/REPRO-Bench.

可复现性AI评估社科研究智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。