用强化学习让论文评审更懂图表和外部信息
REM-CTX: Automated Peer Review via Reinforcement Learning with Auxiliary Context

- 用奖励函数引导模型关注图表和外部文献等辅助信息
- 在多个学科上优于六种基线,甚至超过更大商用模型
- 适合需要提升评审全面性和准确性的研究者
大多数自动化同行评审系统仅依赖文本内容,忽视了图表等视觉元素和外部学术信号。我们提出REM-CTX,一种基于强化学习的评审生成系统,通过对应感知的奖励函数引入辅助上下文。该系统使用80亿参数语言模型,采用分组相对策略优化(GRPO)训练,并融合多维度质量奖励与两项对应奖励,显式鼓励模型与辅助上下文对齐。在计算机、生物和物理科学领域的稿件上实验表明,REM-CTX在六种基线中综合评审质量最高,显著优于其他使用更大商业模型的系统,且在质量和上下文契合度指标上均超越次优的强化学习基线。消融实验确认两项对应奖励具有互补性:各自针对性提升对应奖励表现,同时保持所有质量维度不变,完整模型优于所有简化版本。训练动态分析显示,批评维度与其他指标呈负相关,提示未来研究应针对多维奖励进行分组设计。
原文摘要 · Abstract (English)
Most automated peer review systems rely on textual manuscript content alone, leaving visual elements such as figures and external scholarly signals underutilized. We introduce REM-CTX, a reinforcement-learning system that incorporates auxiliary context into the review generation process via correspondence-aware reward functions. REM-CTX trains an 8B-parameter language model with Group Relative Policy Optimization (GRPO) and combines a multi-aspect quality reward with two correspondence rewards that explicitly encourage alignment with auxiliary context. Experiments on manuscripts across Computer, Biological, and Physical Sciences show that REM-CTX achieves the highest overall review quality among six baselines, outperforming other systems with substantially larger commercial models, and surpassing the next-best RL baseline across both quality and contextual grounding metrics. Ablation studies confirm that the two correspondence rewards are complementary: each selectively improves its targeted correspondence reward while preserving all quality dimensions, and the full model outperforms all partial variants. Analysis of training dynamics reveals that the criticism aspect is negatively correlated with other metrics during training, suggesting that future studies should group multi-dimension rewards for review generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。