arXiv:2605.20668cs.CLcs.AI2026-05被引 2

AI评审比人类更准,但也有盲区,适合作为助手而非替代。

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

论文配图:On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
图 1 · 摘自论文原文
  • 用45位专家人工标注2960条批评,评估AI与人类评审表现。
  • GPT-5.2综合评分超顶尖人类评审(60.0%对48.2%),且发现26%人类未提的问题。
  • AI问题重复率高、有固定弱点,如领域知识有限、不擅长长文档分析。

随着AI能力提升,其在科学同行评审中的应用日益增多,但其能力和可信度仍存争议。现有评估多聚焦于AI与人类意见是否一致,不足以揭示其真实能力边界。本文通过大规模专家标注研究,45位物理、生物与健康科学领域的专家耗时469小时,对82篇Nature系列论文的2960条人类与AI生成的批评进行评分,涵盖正确性、重要性与证据充分性。结果显示,基于GPT-5.2的评审系统在三项综合指标上超越最优秀的人类评审(60.0% vs. 48.2%,p=0.009),所有三款AI(含Gemini 3.0 Pro和Claude Opus 4.5)在每项维度均高于最低人类评分。AI提出的准确批评更常被评价为重要且证据充分,并发现26%人类未提及的问题。但AI评审者之间重合度高达21%(人类仅3%),且存在16种人类未共有的重复弱点,如子领域知识不足、无法处理跨文件长上下文、对小问题过度严苛。总体而言,当前AI评审是人类的补充,而非替代。

原文摘要 · Abstract (English)

With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientists simply view them as probabilistic systems without the expertise to evaluate research, while other researchers are more optimistic about their readiness without concrete evidence. Understanding what AI reviewers do well, where they fall short, and what challenges remain is essential. However, existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits. In this paper, we close this gap through a large-scale expert annotation study, in which 45 domain scientists in Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual criticisms (each targeting one specific aspect of a paper) from human-written and AI-generated reviews of 82 Nature-family papers on correctness, significance, and sufficiency of evidence. On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009), while all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension. AI reviewers' accurate criticisms are also more often rated significant and well-evidenced, and surface a distinct 26% of issues no human raises. However, AI reviewers overlap far more than humans do (21% vs. 3% for cross-reviewer pairs), and exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues. Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers.

AI评审科学评价大模型人类协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。