AI可有效评估经济学论文质量,但需隐藏作者信息防偏见。
Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers
- 用1220篇匿名论文测试4种大模型,文本内容即可区分研究质量。
- Claude和Gemma能准确反映期刊声誉梯度,GPT最擅长识别AI生成内容。
- 若透露作者信息,模型会偏爱顶尖男性学者和名校论文,需谨慎使用。
本研究评估大型语言模型(LLMs)在学术同行评审中辅助评估经济学研究质量的潜力,检验其是否引入系统性偏差。我们对四种模型(GPT-4o、Claude 3.5、Gemma 3、LLaMA 3.3)进行了大规模实验,覆盖超过29,000次对1,220篇匿名论文的评估,这些论文来自110家未被当前大模型训练数据包含的经济学期刊,并包括一组AI生成投稿。结果显示,大模型仅基于文本内容就能稳定区分高质量与低质量研究,生成的质量梯度与既定期刊声誉高度一致。Claude和Gemma在捕捉梯度方面表现优异,而GPT在检测AI生成内容上尤为突出。第二项实验设计了8,910次评估,考察模型是否复制人类在单盲评审中的偏见。通过系统改变330篇论文的作者性别、机构背景和学术声望,发现GPT、Gemma和LLaMA对顶尖男性作者及精英机构的稿件评分显著更高,相比匿名版本。这表明在编辑筛选中部署大模型时,必须去除作者标识信息。总体而言,研究为将大模型整合进同行评审流程提供了有力证据与实践指引,有助于提升效率、准确性与公平性。
原文摘要 · Abstract (English)
This study examines the potential of large language models (LLMs) to augment the academic peer review process by reliably evaluating the quality of economics research without introducing systematic bias. We conduct one of the first large-scale experimental assessments of four LLMs (GPT-4o, Claude 3.5, Gemma 3, and LLaMA 3.3) across two complementary experiments. In the first, we use nonparametric binscatter and linear regression techniques to analyze over 29,000 evaluations of 1,220 anonymized papers drawn from 110 economics journals excluded from the training data of current LLMs, along with a set of AI-generated submissions. The results show that LLMs consistently distinguish between higher- and lower-quality research based solely on textual content, producing quality gradients that closely align with established journal prestige measures. Claude and Gemma perform exceptionally well in capturing these gradients, while GPT excels in detecting AI-generated content. The second experiment comprises 8,910 evaluations designed to assess whether LLMs replicate human like biases in single blind reviews. By systematically varying author gender, institutional affiliation, and academic prominence across 330 papers, we find that GPT, Gemma, and LLaMA assign significantly higher ratings to submissions from top male authors and elite institutions relative to the same papers presented anonymously. These results emphasize the importance of excluding author-identifying information when deploying LLMs in editorial screening. Overall, our findings provide compelling evidence and practical guidance for integrating LLMs into peer review to enhance efficiency, improve accuracy, and promote equity in the publication process of economics research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。