arXiv:2604.13940cs.AI2026-04综述被引 16

AI为2.3万篇论文生成审查意见,效果优于人工。

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

论文配图:AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
图 1 · 摘自论文原文
  • 用先进AI系统多阶段生成评审,一天内完成全部论文
  • 22,977篇论文的评审中,AI意见在技术准确性和建议上更受好评
  • 首次大规模实证展示AI可胜任会议级评审任务

科学同行评审面临投稿量激增带来的压力,难以维持质量、一致性和时效性。近期人工智能的发展促使学界思考其在评审中的应用,但关键问题在于:AI能否在真实会议规模下生成技术上可靠的评审意见?本文报告了首次大规模实地部署的AI辅助同行评审:AAAI-26所有主赛道投稿均获得一份明确标识的AI生成评审。该系统结合前沿模型、工具调用与安全机制,以多阶段流程在不到一天时间内完成了对22,977篇完整论文的评审。对作者和程序委员会成员的大规模调查显示,参与者不仅认为AI评审有用,甚至在技术准确性与研究建议等关键维度上更偏好AI评审。我们还引入了一个新基准,发现该系统在检测多种科学缺陷方面显著优于简单的LLM生成基线。这些结果表明,当前最先进的AI方法已能在会议规模上对科学评审做出实质性贡献,开启了人机协同评估研究的新路径。

原文摘要 · Abstract (English)

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer review, yet a key unresolved question is whether AI can generate technically sound reviews at real-world conference scale. Here we report the first large-scale field deployment of AI-assisted peer review: every main-track submission at AAAI-26 received one clearly identified AI review from a state-of-the-art system. The system combined frontier models, tool use, and safeguards in a multi-stage process to generate reviews for all 22,977 full-review papers in less than a day. A large-scale survey of AAAI-26 authors and program committee members showed that participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions. We also introduce a novel benchmark and find that our system substantially outperforms a simple LLM-generated review baseline at detecting a variety of scientific weaknesses. Together, these results show that state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research.

同行评审AI辅助大规模实验学术智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。