评测大模型写论文评审的行为,发现它和人有明显差异。
PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
- 构建新基准PRAIB,量化大模型评审的细节、风格与互动行为
- 大模型评审更长更复杂,但评分偏差大且常漏掉关键问题
- 适合评估审稿辅助工具,不建议直接替代人类审稿
随着投稿量增加,大语言模型(LLMs)被探索用于提升同行评审的速度与可扩展性。然而,尚不清楚大模型是否像人类审稿人一样真正理解论文,还是仅生成看似合理的评审文本。为此,我们提出同行评审人工智能基准(PRAIB),包含衡量评审具体性、风格与互动行为的多维指标。基于1,000篇ICLR与NeurIPS论文(2021–2025年)生成的11,000条评审数据,涵盖五种专有与开源模型,在多种提示策略下进行大规模实证研究。结果表明,机器生成评审在评分上变异性低、存在正向偏见且过度自信,引用模式也依赖模型且不同于人类习惯。尽管生成文本更长更复杂,却常忽略人类审稿人指出的细粒度缺陷。PRAIB为社区提供了诊断工具,明确当前大模型可可靠支持的审稿环节及其仍需改进之处。
原文摘要 · Abstract (English)
The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review process, particularly in terms of improving its speed and scalability. Yet, it remains unknown whether LLMs engage with scientific manuscripts in the same manner as human reviewers, or whether they merely produce review-looking text. To address this, we introduce the Peer Review AI Benchmark (PRAIB), a novel framework comprising thoroughly defined metrics that measure review specificity, style, and behavior of engagement. To complement the PRAIB framework, we conduct a large-scale empirical study leveraging a dataset of 11,000 reviews generated by five proprietary and open-source models for 1,000 ICLR and NeurIPS papers. Spanning the 2021--2025 period, these machine-generated reviews are compared against original human feedback across diverse prompting strategies to identify systematic behavioral divergences. Our analysis reveals that the generated reviews diverge significantly from feedback provided by human reviewers: LLM ratings are less variable, positively biased, and overconfident, and their cross-reference patterns are model-dependent and distinct from human norms. Furthermore, when evaluated through PRAIB, we observe that LLMs tend to generate longer, more complex reviews, yet frequently overlook the atomic weaknesses noted by human reviewers. By characterizing where and how LLMs reviewing behavior departs from human norms, PRAIB provides the community with a diagnostic tool for identifying which aspects of the review process LLMs can reliably support today and which require further development before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。