构建首个兼顾完整性和正确性的AI审稿评估基准。
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

- 按类别构建子集,缺失人类评审时跳过评估以提升完整性。
- 利用审稿人-作者-元评审对话作为专家标注,过滤不可靠评论。
- 基于ICLR和NeurIPS的3900篇论文,揭示当前AI审稿存在幻觉问题。
尽管AI审稿系统发展迅速,但其评估仍具挑战性:现有指标偏向与人类审稿的重合度,而非正确性。然而人类审稿常遗漏关键问题或包含错误,无法作为可靠参考。为此,我们构建了类别特定的基准子集,在对应人类审稿缺失时跳过评估以增强完整性;同时利用审稿人-作者-元评审讨论作为专家标注,筛选出不可靠的审稿意见以提升正确性。最终提出CoCoReviewBench,收录ICLR与NeurIPS的3,900篇论文,支持对AI审稿系统的可靠、细粒度评估。分析显示,当前AI审稿在正确性方面仍有局限,易产生幻觉,且推理型模型表现更优,为未来改进提供方向。代码与模型已开源。
原文摘要 · Abstract (English)
Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews often cover only a subset of salient issues and sometimes contain mistakes, they are unreliable as gold references. To address this, we build category-specific benchmark subsets and skip evaluation when the corresponding human reviews are missing to strengthen Completeness. We also leverage reviewer--author--meta-review discussions as expert annotations and filter unreliable reviews accordingly to strengthen Correctness. Finally, we introduce CoCoReviewBench, which curates 3,900 papers from ICLR and NeurIPS to enable reliable and fine-grained evaluation of AI reviewers. Analysis shows that AI reviewers remain limited in correctness and are prone to hallucinations, and highlights reasoning models as more effective reviewers, motivating further directions for improving AI reviewers. Benchmarks and models are available at https://github.com/hexuandeng/CoCoReviewBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。