让AI评审报告可追溯,每条意见都有证据和操作建议。
DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review

- 构建可追踪的评审清单,自动关联论文内容与证据
- 在ICLR 2025上比Gemini表现更好,重大问题覆盖率达37.26%
- 适合需要透明、可审计评审过程的研究团队使用
自动化同行评审常被理解为生成流畅批评,但审稿人和领域主席需要可审计的判断:问题出现在何处、依据是什么、应如何跟进。DeepReviewer 2.0 是一个基于输出契约的过程控制型智能体系统,生成包含锚定标注、局部证据和可执行后续动作的可追溯评审包,并仅在满足最低可追溯性与覆盖率预算后才导出结果。具体而言,它先构建仅基于论文的主张-证据-风险清单与验证议程,再按议程驱动检索并撰写带锚点的批判性意见,通过出口门控机制确保质量。在134篇ICLR 2025投稿中,未微调的1960亿参数模型在三种固定协议下表现优于Gemini-3.1-Pro-preview,严格重大问题覆盖率达37.26%(对比23.57%),在盲评中以71.63%的微平均胜率击败人类评审委员会,成为本组自动系统排名第一。我们定位其为辅助工具而非决策代理,并指出伦理敏感性检查等仍待完善之处。
原文摘要 · Abstract (English)
Automated peer review is often framed as generating fluent critique, yet reviewers and area chairs need judgments they can \emph{audit}: where a concern applies, what evidence supports it, and what concrete follow-up is required. DeepReviewer~2.0 is a process-controlled agentic review system built around an output contract: it produces a \textbf{traceable review package} with anchored annotations, localized evidence, and executable follow-up actions, and it exports only after meeting minimum traceability and coverage budgets. Concretely, it first builds a manuscript-only claim--evidence--risk ledger and verification agenda, then performs agenda-driven retrieval and writes anchored critiques under an export gate. On 134 ICLR~2025 submissions under three fixed protocols, an \emph{un-finetuned 196B} model running DeepReviewer~2.0 outperforms Gemini-3.1-Pro-preview, improving strict major-issue coverage (37.26\% vs.\ 23.57\%) and winning 71.63\% of micro-averaged blind comparisons against a human review committee, while ranking first among automatic systems in our pool. We position DeepReviewer~2.0 as an assistive tool rather than a decision proxy, and note remaining gaps such as ethics-sensitive checks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。