构建首个大规模AI与人工撰写审稿意见数据集,助力科研评审可信性研究。
Gen-Review: A Large-scale Dataset of AI-Generated (and Human-written) Peer Reviews
- 用三种提示生成8.1万条AI审稿意见,覆盖ICLR 2018–2025所有投稿
- 发现AI审稿存在偏见、可被检测、指令遵循不严格,仅通过录用率的评分较可靠
- 支持多维度研究,适合关注AI伦理、审稿机制、内容检测的研究者
大型语言模型(LLMs)对科学同行评审的影响日益显著,但相关数据匮乏。本文提出GenReview,迄今最大的包含AI生成与人工撰写审稿意见的数据集,涵盖ICLR 2018–2025年全部投稿的81,000条审稿意见,每篇论文由LLM基于负面、正面和中立三种提示生成。该数据集关联原始论文及真实审稿意见,支持多种研究。我们验证了若干关键问题:AI审稿存在偏见;目前可被自动识别;对指令遵循不一致;仅在已录用论文上评分与接受决策高度一致。数据集可通过https://anonymous.4open.science/r/gen_review获取。
原文摘要 · Abstract (English)
How does the progressive embracement of Large Language Models (LLMs) affect scientific peer reviewing? This multifaceted question is fundamental to the effectiveness -- as well as to the integrity -- of the scientific process. Recent evidence suggests that LLMs may have already been tacitly used in peer reviewing, e.g., at the 2024 International Conference of Learning Representations (ICLR). Furthermore, some efforts have been undertaken in an attempt to explicitly integrate LLMs in peer reviewing by various editorial boards (including that of ICLR'25). To fully understand the utility and the implications of LLMs' deployment for scientific reviewing, a comprehensive relevant dataset is strongly desirable. Despite some previous research on this topic, such dataset has been lacking so far. We fill in this gap by presenting GenReview, the hitherto largest dataset containing LLM-written reviews. Our dataset includes 81K reviews generated for all submissions to the 2018--2025 editions of the ICLR by providing the LLM with three independent prompts: a negative, a positive, and a neutral one. GenReview is also linked to the respective papers and their original reviews, thereby enabling a broad range of investigations. To illustrate the value of GenReview, we explore a sample of intriguing research questions, namely: if LLMs exhibit bias in reviewing (they do); if LLM-written reviews can be automatically detected (so far, they can); if LLMs can rigorously follow reviewing instructions (not always) and whether LLM-provided ratings align with decisions on paper acceptance or rejection (holds true only for accepted papers). GenReview can be accessed at the following link: https://anonymous.4open.science/r/gen_review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。