arXiv:2510.08867cs.AIcs.CL2025-10综述被引 11

AI可辅助审稿,提升一致性与覆盖度,但复杂判断仍需人类专家。

ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review

  • 构建模块化框架ReviewerToo,支持模拟不同审稿人角色和标准评估。
  • 在ICLR 2025 1963篇论文上,AI审稿准确率达81.8%,接近人类平均的83.9%。
  • 适合关注审稿公平性、可扩展性的科研管理者与会议组织者。

同行评审是科学出版的核心,但存在不一致、主观性强及可扩展性差等问题。我们提出ReviewerToo,一个用于研究和部署人工智能辅助同行评审的模块化框架,旨在以系统化、一致性的评估补充人类判断。该框架支持针对特定审稿人角色和结构化评价标准的系统性实验,并可部分或完全融入真实会议流程。我们在精心整理的ICLR 2025 1,963篇投稿数据集上验证了该框架,使用gpt-oss-120b模型在判别论文是否接收的任务中达到81.8%的准确率,略低于人类平均的83.9%。此外,由ReviewerToo生成的评审意见被大语言模型评委评为质量高于人类平均水平,但仍不及最强专家贡献。分析显示,AI在事实核查、文献覆盖方面表现优异,但在方法新颖性和理论贡献评估上仍显不足,凸显人类专业知识的必要性。基于此,我们提出集成AI的指导原则,表明AI可提升评审的一致性、覆盖面与公平性,而复杂判断仍交由领域专家处理。本工作为可扩展的混合式同行评审体系奠定基础。

原文摘要 · Abstract (English)

Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular framework for studying and deploying AI-assisted peer review to complement human judgment with systematic and consistent assessments. ReviewerToo supports systematic experiments with specialized reviewer personas and structured evaluation criteria, and can be partially or fully integrated into real conference workflows. We validate ReviewerToo on a carefully curated dataset of 1,963 paper submissions from ICLR 2025, where our experiments with the gpt-oss-120b model achieves 81.8% accuracy for the task of categorizing a paper as accept/reject compared to 83.9% for the average human reviewer. Additionally, ReviewerToo-generated reviews are rated as higher quality than the human average by an LLM judge, though still trailing the strongest expert contributions. Our analysis highlights domains where AI reviewers excel (e.g., fact-checking, literature coverage) and where they struggle (e.g., assessing methodological novelty and theoretical contributions), underscoring the continued need for human expertise. Based on these findings, we propose guidelines for integrating AI into peer-review pipelines, showing how AI can enhance consistency, coverage, and fairness while leaving complex evaluative judgments to domain experts. Our work provides a foundation for systematic, hybrid peer-review systems that scale with the growth of scientific publishing.

同行评审AI辅助学术出版评分系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。