arXiv:2505.07920cs.CLcs.AI2025-05综述被引 19

构建首个全流程同行评审数据集,支持多轮反驳对话。

Re2: A Consistency-ensured Dataset for Full-stage Peer Review and Multi-turn Rebuttal Discussions

  • 基于原始投稿构建一致性保障数据集,避免修改后内容干扰
  • 涵盖19,926篇初始投稿、7万余条评审意见与5万余条反驳文本
  • 支持多轮交互式对话,适用于AI辅助审稿与作者自评

同行评审是人工智能等领域科学进步的关键环节,但投稿量激增导致评审系统超载,引发评审人力不足与质量下降。其中一个关键原因在于大量低质量稿件反复提交,根源在于缺乏有效的投稿前自我评估工具。大语言模型(LLMs)在辅助作者与审稿人方面潜力巨大,但其性能受限于同行评审数据的质量。现有数据集存在三大缺陷:(1) 数据多样性不足;(2) 因使用修改后投稿导致内容不一致且质量偏低;(3) 缺乏对反驳与互动讨论任务的支持。为此,我们提出Re²——目前最大且具一致性保障的同行评审与反驳数据集,包含来自24个会议和21个研讨会的19,926篇初始投稿、70,668条评审意见与53,818条反驳内容。该数据集将反驳与讨论阶段建模为多轮对话范式,既支持传统静态评审任务,也适配动态交互式大模型助手,为作者优化论文提供更实用指导,并缓解日益增长的评审负担。数据与代码已公开于https://anonymous.4open.science/r/ReviewBench_anon/。

原文摘要 · Abstract (English)

Peer review is a critical component of scientific progress in the fields like AI, but the rapid increase in submission volume has strained the reviewing system, which inevitably leads to reviewer shortages and declines review quality. Besides the growing research popularity, another key factor in this overload is the repeated resubmission of substandard manuscripts, largely due to the lack of effective tools for authors to self-evaluate their work before submission. Large Language Models (LLMs) show great promise in assisting both authors and reviewers, and their performance is fundamentally limited by the quality of the peer review data. However, existing peer review datasets face three major limitations: (1) limited data diversity, (2) inconsistent and low-quality data due to the use of revised rather than initial submissions, and (3) insufficient support for tasks involving rebuttal and reviewer-author interactions. To address these challenges, we introduce the largest consistency-ensured peer review and rebuttal dataset named Re^2, which comprises 19,926 initial submissions, 70,668 review comments, and 53,818 rebuttals from 24 conferences and 21 workshops on OpenReview. Moreover, the rebuttal and discussion stage is framed as a multi-turn conversation paradigm to support both traditional static review tasks and dynamic interactive LLM assistants, providing more practical guidance for authors to refine their manuscripts and helping alleviate the growing review burden. Our data and code are available in https://anonymous.4open.science/r/ReviewBench_anon/.

同行评审多轮对话数据集LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。