用强化学习训练论文评审与辩护智能体,提升推理深度与引用准确性。
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

- 基于统一客观奖励机制的智能体强化学习框架
- 引用准确率显著提升,有效杜绝幻觉现象
- 适合需要高质量学术写作的科研人员与AI开发者
生成专业的学术内容,如同行评审和反驳意见,需要领域推理与事实依据的紧密结合。本文提出一个完整的框架,用于开发与评估专门的学术智能体——InternReviewer 和 InternAdvocate。首先构建了一个大规模、高质量的学术数据集,并集成高效的 arXiv 检索工具以支持主动证据获取。为优化这些智能体,我们采用由统一目标指标驱动的智能体强化学习(agentic RL)范式。该系统通过多维标准避免主观模型评判带来的偏差,包括参考锚定的语义对齐、结构合规性以及严格的验证机制——交叉核对引文与实时交互日志,消除幻觉。实验结果表明,在闭环框架中训练的智能体在推理深度和引用准确性方面均有显著提升。
原文摘要 · Abstract (English)
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。