arXiv:2503.10671cs.CL2025-03被引 2

用大模型判断社科研究能否复现,效果接近人类水平。

Identifying Non-Replicable Social Science Studies with Language Models

  • 用大模型生成行为研究模拟数据,检验原结论是否成立。
  • 最佳模型(Mistral 7B)F1达77%,接近人类复现结果。
  • 低采样温度会引入偏差,影响效应量估计准确性。

本研究探讨大语言模型(LLM)能否识别行为社会科学领域研究的可复现性。基于14项已复现研究的数据集(9次成功,5次失败),评估了开源模型(Llama 3 8B、Qwen 2 7B、Mistral 7B)和专有模型(GPT-4o)在区分可复现与不可复现发现上的能力。通过让模型生成行为研究的合成响应样本,评估测量效应是否支持原结论。与人类复现结果对比,Mistral 7B达到最高77%的F1值,GPT-4o和Llama 3 8B均为67%,Qwen 2 7B为55%。此外,分析发现采样温度过低会导致方差减小,从而产生效应量估计偏差。

原文摘要 · Abstract (English)

In this study, we investigate whether LLMs can be used to indicate if a study in the behavioural social sciences is replicable. Using a dataset of 14 previously replicated studies (9 successful, 5 unsuccessful), we evaluate the ability of both open-source (Llama 3 8B, Qwen 2 7B, Mistral 7B) and proprietary (GPT-4o) instruction-tuned LLMs to discriminate between replicable and non-replicable findings. We use LLMs to generate synthetic samples of responses from behavioural studies and estimate whether the measured effects support the original findings. When compared with human replication results for these studies, we achieve F1 values of up to $77\%$ with Mistral 7B, $67\%$ with GPT-4o and Llama 3 8B, and $55\%$ with Qwen 2 7B, suggesting their potential for this task. We also analyse how effect size calculations are affected by sampling temperature and find that low variance (due to temperature) leads to biased effect estimates.

大模型应用可复现性社科研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。