arXiv:2503.08506cs.CL2025-03被引 30

用AI生成更像人写的论文评审,提升自动化质量。

ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews

  • 构建多智能体框架,模仿人类评审的结构化思考过程。
  • 在14.2万条评论数据上训练,生成评论更准确、逻辑更连贯。
  • 适合研究者、期刊编辑用于提升审稿效率与一致性。

学术论文评审是科研界关键但耗时的任务。随着论文数量激增,自动化评审面临挑战:如何生成全面、准确且推理一致的评审意见。本文提出ReviewAgents框架,利用大语言模型(LLMs)生成论文评审。首先构建包含14.2万条评审评论的Review-CoT数据集,模拟人类评审的结构化思维流程——总结论文、引用相关工作、识别优缺点并得出结论。基于此,采用关联文献感知训练方法训练出具备结构化推理能力的LLM评审代理。进一步构建ReviewAgents,一个由多个角色、多模型组成的智能体评审系统,以增强评审意见生成。同时提出ReviewBench基准,用于评估生成评论的质量。实验表明,尽管现有LLMs具有一定潜力,但仍不及人工评审;而本框架显著缩小差距,在生成评审意见方面优于先进LLMs。

原文摘要 · Abstract (English)

Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has become a significant challenge. The primary issue lies in generating comprehensive, accurate, and reasoning-consistent review comments that align with human reviewers' judgments. In this paper, we address this challenge by proposing ReviewAgents, a framework that leverages large language models (LLMs) to generate academic paper reviews. We first introduce a novel dataset, Review-CoT, consisting of 142k review comments, designed for training LLM agents. This dataset emulates the structured reasoning process of human reviewers-summarizing the paper, referencing relevant works, identifying strengths and weaknesses, and generating a review conclusion. Building upon this, we train LLM reviewer agents capable of structured reasoning using a relevant-paper-aware training method. Furthermore, we construct ReviewAgents, a multi-role, multi-LLM agent review framework, to enhance the review comment generation process. Additionally, we propose ReviewBench, a benchmark for evaluating the review comments generated by LLMs. Our experimental results on ReviewBench demonstrate that while existing LLMs exhibit a certain degree of potential for automating the review process, there remains a gap when compared to human-generated reviews. Moreover, our ReviewAgents framework further narrows this gap, outperforming advanced LLMs in generating review comments.

AI审稿大模型应用多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。