测试大模型当评审是否公平,发现它偏爱自己写的文章。
LLM-REVal: Can We Trust LLM Reviewers Yet?
- 用模拟研究与评审双代理对比人类与大模型评分差异
- 大模型对自产论文打分高,对人类论文常因批判性表述压分
- 揭示语言风格偏好与排斥批评的双重偏差,警示评审滥用风险
大语言模型(LLMs)快速发展的背景下,其在学术研究与同行评审中的深度集成正重塑科研流程。本文通过构建研究代理(生成论文并修订)与评审代理(评估投稿)的仿真系统,探究大模型作为评审者的潜在风险。结果表明:(1)大模型评审者系统性高估自身生成的论文,得分显著高于人类作者作品;(2)即便经过多次修改,大模型仍持续低估人类作者论文,尤其针对包含“风险”“公平性”等批判性表述的内容。分析显示,根源在于大模型对语言风格的偏好及对批判性内容的回避。尽管如此,基于大模型评审建议的修改可提升论文质量,在人类与大模型评价中均表现更好,表明其对早期研究者和低质量稿件具有改进潜力。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studies highlight the potential of LLMs in supporting research and peer review, their dual roles in the academic workflow and the complex interplay between research and review bring new risks that remain largely underexplored. In this study, we focus on how the deep integration of LLMs into both peer-review and research processes may influence scholarly fairness, examining the potential risks of using LLMs as reviewers by simulation. This simulation incorporates a research agent, which generates papers and revises, alongside a review agent, which assesses the submissions. Based on the simulation results, we conduct human annotations and identify pronounced misalignment between LLM-based reviews and human judgments: (1) LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones; (2) LLM reviewers persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions. Our analysis reveals that these stem from two primary biases in LLM reviewers: a linguistic feature bias favoring LLM-generated writing styles, and an aversion toward critical statements. These results highlight the risks and equity concerns posed to human authors and academic research if LLMs are deployed in the peer review cycle without adequate caution. On the other hand, revisions guided by LLM reviews yield quality gains in both LLM-based and human evaluations, illustrating the potential of the LLMs-as-reviewers for early-stage researchers and enhancing low-quality papers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。