用强化学习生成更准确、深入的学术论文评审,解决自动评审泛化问题。
ReviewRL: Towards Automated Scientific Review with RL
- 结合文献检索与监督微调,构建可理解科学内容的评审模型。
- 在ICLR 2025论文上表现优于现有方法,评分更准、反馈更深入。
- 适合科研机构和期刊用于辅助审稿,提升评审效率与一致性。
同行评审是科学进步的关键,但面对投稿量激增和审稿人疲劳,现有自动化评审方法在事实准确性、评分一致性和分析深度上仍显不足,常生成浅显或通用的反馈。我们提出ReviewRL,一种基于强化学习的科学论文评审生成框架。该方法融合:(1)基于ArXiv-MCP的检索增强上下文生成管道,引入相关文献;(2)监督微调以建立基础评审能力;(3)复合奖励函数驱动的强化学习过程,联合优化评审质量与评分准确性。在ICLR 2025论文上的实验表明,ReviewRL显著优于现有方法,无论在规则指标还是模型评估中均表现优异。该工作为科学发现中的强化学习驱动自动评述奠定了基础,展现出广阔的应用前景。代码将发布于GitHub。
原文摘要 · Abstract (English)
Peer review is essential for scientific progress but faces growing challenges due to increasing submission volumes and reviewer fatigue. Existing automated review approaches struggle with factual accuracy, rating consistency, and analytical depth, often generating superficial or generic feedback lacking the insights characteristic of high-quality human reviews. We introduce ReviewRL, a reinforcement learning framework for generating comprehensive and factually grounded scientific paper reviews. Our approach combines: (1) an ArXiv-MCP retrieval-augmented context generation pipeline that incorporates relevant scientific literature, (2) supervised fine-tuning that establishes foundational reviewing capabilities, and (3) a reinforcement learning procedure with a composite reward function that jointly enhances review quality and rating accuracy. Experiments on ICLR 2025 papers demonstrate that ReviewRL significantly outperforms existing methods across both rule-based metrics and model-based quality assessments. ReviewRL establishes a foundational framework for RL-driven automatic critique generation in scientific discovery, demonstrating promising potential for future development in this domain. The implementation of ReviewRL will be released at GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。