arXiv:2606.19749cs.AIcs.CL2026-06综述

评测AI评审系统在真实论文上的表现,发现其能有效识别错误并贴近人类评价。

Benchmarking Agentic Review Systems

论文配图:Benchmarking Agentic Review Systems
图 1 · 摘自论文原文
  • 用六种大模型测试三种AI评审系统,比较其判断论文质量的能力
  • 最佳系统(OpenAIReview+GPT-5.5)准确率达83.0%,检测错误召回率71.6%
  • 多模型协作可提升至83.3%召回率,适合研究者与审稿人参考

针对人工智能辅助研究对同行评审带来的压力,新兴的智能评审系统亟需评估方法。本文评估了两个开源系统(OpenAIReview和coarse)、一个专有系统(Reviewer3)及零样本基线,在六种大语言模型(涵盖前沿与高效模型)上进行测试。首先,通过引用次数和录用结果等外部信号衡量论文质量,发现所有系统在成对判断中均优于随机水平,最佳组合(OpenAIReview + GPT-5.5)达到83.0%的准确率。其次,构建扰动基准,向八个arXiv领域论文注入四类错误,测量系统检测召回率。最强配置检出71.6%错误,仍有提升空间;六模型联合检测达83.3%召回率,表明不同模型捕捉不同错误,协同设计可进一步提升性能。此外,公开部署的OpenAIReview获用户反馈,评论评分正向偏倚为1.44:1,主要批评集中在误报和琐碎问题。综合表明,尽管尚有改进空间,当前AI评审系统已能较好追踪人类质量判断、发现关键错误,并获得真实用户认可。

原文摘要 · Abstract (English)

A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but it is unclear how they should be evaluated. We evaluate two open-source systems (OpenAIReview and coarse), one proprietary system (Reviewer3), and a zero-shot baseline, across six LLMs spanning frontier and efficient models. First, we study whether AI reviews on ICLR/NeurIPS papers track with papers' quality as approximated by external signals such as citations and acceptance decisions. Every system performs above chance in pairwise accuracy, and the best is OpenAIReview + GPT-5.5 at 83.0%. Second, to test whether systems can catch errors with known ground truth, we construct a perturbation benchmark that injects four categories of errors into papers across eight arXiv subject classes and measure detection recall. The strongest configuration (OpenAIReview + GPT-5.5) catches 71.6% of injected errors, leaving substantial room for improvement. The union of detections across six models reaches 83.3% recall, suggesting different models detect different errors and better harness design can potentially increase performance. Beyond these benchmarks, we study a public deployment of OpenAIReview with real users. Votes on its comments skew positive at 1.44 to 1, and the most common complaints are about false positives and minor nitpicks. Together, by evaluating full review systems backed by state-of-the-art models on real research papers, we show that while AI reviews still have room for improvement, they can already track human quality judgments well, catch important errors, and earn positive feedback from real users.

AI评审大模型评测基准论文质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。