用大模型自动评估社科论文可复现性,效率远超人工。
Automated reproducibility assessments in the social and behavioral sciences using large language models

- 用大模型自动重分析180篇社科论文的原始数据。
- 80%情况下结论与原研究一致,24%恢复了相同效应量。
- 适合用于大规模筛查,辅助科研严谨性审核。
社会科学中的可复现性通常由独立研究人员重新分析原始数据来评估,但该方法耗时且难规模化。本文展示大语言模型(LLMs)可自动化完成可复现性评估。基于180篇行为与社会科学领域发表的研究,其预设结论被用于对比。在11项研究中,大模型无法生成有效效应量估计;其余研究中,大模型在80%的情况下得出与原研究相同的定性结论,且在24%的研究中恢复了原始效应量(以Cohen's d ±0.05为容差)。在包含人工重分析的子集里,大模型在95%的研究中达成相同定性结论,接近人类重分析者(83%),并以40%的准确率恢复原始效应量,与人类(28%)相当。当前大模型能力虽有限,但足以作为系统性审计工具,提升实证研究的严谨性与可复现性。
原文摘要 · Abstract (English)
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. However, such approaches are resource-intensive and difficult to scale. Here, we show that large language models (LLMs) can automate reproducibility assessments. Using N = 180 published studies with predefined claims from the behavioral and social sciences, we compare LLM-generated analyses with the original findings. For 11 studies, the LLM pipeline could not produce a viable effect size estimate. For the remaining studies, the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a +/-0.05 tolerance in Cohen's d) in 24% of studies. In a subset with human reanalyses, the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a +/-0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%). Given the current capabilities and limitations of LLMs, the findings show that LLMs can support systematic audits of empirical results rather than substitute expert judgment. As such, LLMs can serve as a scalable screening tool to improve the rigor and reproducibility in empirical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。