让AI像真人一样主动查证论文疑点,提升审稿质量
From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

- 用结构化审查日志构建可追踪的推理过程
- 8B模型经强化学习优化后评分比大模型高39%
- 适合需要高质量自动化审稿的科研机构使用
大型语言模型在自动化科学同行评审方面展现出潜力,但现有方法常难以生成基于具体证据的深入评审。我们认为关键瓶颈在于缺乏主动调查论文可疑部分的能力,而人类审稿人恰恰具备这种能力。本文探索如何让基于LLM的评审代理实现主动调查,将该问题自然建模为马尔可夫决策过程,并提出ProReviewer:一个基于结构化审查日志进行主动评审的科学同行评审代理。该日志作为工作空间,帮助代理追踪评审过程中收集的证据与中间发现。实验表明,采用8B主干模型、通过监督微调训练并经强化学习优化的ProReviewer,在五个质量维度上的平均得分最高,相比使用更大前沿模型的提示工程方法最高提升39%,相对于最强微调基线相对提升16%。在人工评估中,其胜率也高于各基线。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in automating scientific peer review. However, existing approaches often struggle to generate in-depth reviews supported by concrete evidence. We argue that a key limitation is the lack of flexibility to proactively investigate suspicious parts of a paper based on accumulated evidence, as human reviewers do. In this paper, we explore how to enable an LLM-based review agent to perform such proactive investigation. We find that this can be naturally formulated as a Markov Decision Process (MDP), and propose ProReviewer, a scientific peer review agent that proactively reviews a paper guided by a maintained, structured review log. The structured review log serves as a workspace for the agent to track evidence and intermediate findings collected during review. Experiments show that ProReviewer with an 8B backbone, trained by supervised fine-tuning and optimized by reinforcement learning, achieves the highest average score across five quality dimensions, outperforming prompt-based methods with much larger frontier LLMs by up to 39% and the strongest fine-tuned baseline by 16% relatively. It also attains the highest win rates against baselines in human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。