AI评阅需避免过度批评,新方法通过证据修正缺陷,提升审查准确性。
More Criticism Does Not Make a Better Review: EquiReview-R

- 将AI评阅重构为基于证据的关切点修正,区分遗漏与过度批评风险
- 在未见论文上减少重大误判率至8.1%,遗漏上限为9.9%,52.4%论文可停止评审
- 提出新框架EquiReview-R,支持停止、继续或延迟决策,适合严谨学术评审场景
AI评阅虽能生成大量具体批评,但更多批评不等于更好评审。评阅可能遗漏关键缺陷,或保留缺乏证据支持的指控。这两种错误需相反修正,但传统生成系统与聚合指标掩盖了这一区别。为此,我们将AI辅助评审重新定义为基于证据的结构化关切点精炼过程,将遗漏和过度批评视为独立风险。在此基础上,我们提出EquiReview-R:它针对局部证据解决现有关切,从独立且受评审条件影响的角度搜寻缺失问题,并返回停止、继续或推迟决策。为揭示设计动机,我们构建了证据关联轨迹语料库(ReviewTrace)。其回顾性分析表明,修订必须先于进一步搜索:高召回评审中的几乎全部关切均无明确证据结论,而早期修正机制无法处理这些关切。在冻结的未见论文集上,EquiReview-R满足重大遗漏的预设非劣性标准,将重大过度批评从15.5%降至8.1%,并实现9.9%的单侧遗漏上限,同时在52.4%的论文上停止评审。计算匹配对照组、配对实验与消融测试表明,性能提升源于修订机制,而非额外推理或更短输出。我们公开发布语料库ReviewTrace,作为研究评审修订、分歧与溯源的证据关联资源。
原文摘要 · Abstract (English)
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。