用大模型模拟审稿人,自动给出专业、透明的论文评审意见。
Generative Adversarial Reviews: When LLMs Become the Critic
- 构建带记忆和人格化的智能审稿代理,基于论文图结构分析内容逻辑。
- 在真实论文数据集上,评审质量与人类审稿人相当,接受率预测准确率超85%。
- 适合科研人员早期获取专家反馈,尤其利于青年学者和小团队使用。
同行评审是科学进步的核心,决定哪些论文达到出版标准。然而,学术产出激增与知识领域日益细分,使传统反馈机制面临压力。为此,我们提出生成式代理审稿人(Generative Agent Reviewers, GAR),利用大语言模型驱动的智能体模拟忠实的同行审稿人。为实现生成式审稿,我们设计了具备记忆能力的模型架构,并从历史数据中提取审稿人格特征赋予代理。核心是采用图结构表示论文,将观点与证据、技术细节进行逻辑关联。GAR通过外部知识评估论文新颖性,再基于图结构进行多轮详细评估,最后由元审稿人聚合意见预测录用结果。实验表明,GAR在提供详细反馈和预测论文结果方面表现接近人类审稿人。我们还开展深入实验,如评估审稿人专业度影响及评审公平性。GAR可为原本仅限少数研究者获得的专家级反馈提供早期、透明、可及的评估,推动评审过程民主化。
原文摘要 · Abstract (English)
The peer review process is fundamental to scientific progress, determining which papers meet the quality standards for publication. Yet, the rapid growth of scholarly production and increasing specialization in knowledge areas strain traditional scientific feedback mechanisms. In light of this, we introduce Generative Agent Reviewers (GAR), leveraging LLM-empowered agents to simulate faithful peer reviewers. To enable generative reviewers, we design an architecture that extends a large language model with memory capabilities and equips agents with reviewer personas derived from historical data. Central to this approach is a graph-based representation of manuscripts, condensing content and logically organizing information - linking ideas with evidence and technical details. GAR's review process leverages external knowledge to evaluate paper novelty, followed by detailed assessment using the graph representation and multi-round assessment. Finally, a meta-reviewer aggregates individual reviews to predict the acceptance decision. Our experiments demonstrate that GAR performs comparably to human reviewers in providing detailed feedback and predicting paper outcomes. Beyond mere performance comparison, we conduct insightful experiments, such as evaluating the impact of reviewer expertise and examining fairness in reviews. By offering early expert-level feedback, typically restricted to a limited group of researchers, GAR democratizes access to transparent and in-depth evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。