arXiv:2606.28277cs.LGcs.AI2026-06综述

用AI工具自动审查论文,提升科研评审效率。

Towards Automating Scientific Review with Google's Paper Assistant Tool

论文配图:Towards Automating Scientific Review with Google's Paper Assistant Tool
图 1 · 摘自论文原文
  • 构建智能代理框架,深度分析论文全内容
  • 在数学错误识别上比零样本提升34%准确率
  • 适合投稿前自查的作者与审稿人参考

人工智能正加速科学发现,从假说生成到定理证明。然而,传统人工同行评审难以应对由此带来的海量论文压力。为缓解这一矛盾,需用AI加速验证与评审流程。我们提出一个四级渐进式人机协作评审分类体系,分析各阶段权衡。作为推进举措,我们推出论文助手工具(Paper Assistant Tool, PAT),一种面向深度科学评审的智能代理框架。PAT可接收完整论文,评估理论结果、验证实验、建议改进并识别潜在缺陷。通过推理扩展技术,其在SPOT基准上的数学错误召回率比零样本提升34%。在STOC和ICML两个顶级计算机会议的试点中,PAT成功发现关键错误并提出实质性改进建议。早期纠错减轻了审稿人认知负担,同时保留其最终决策权。

原文摘要 · Abstract (English)

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving. However, this rapid acceleration is creating a systemic challenge: traditional human peer review cannot scale to match the influx of AI-assisted science. Ultimately, to resolve this tension, we must also deploy AI to accelerate the verification and review process itself. To frame the discussion around this transition, we propose a taxonomy consisting of four progressive levels of AI-human collaboration in scientific evaluation, and discuss various trade-offs involved with each. As a step toward this future, we introduce the Paper Assistant Tool (PAT), an agentic AI framework built for deep scientific review and verification. PAT ingests full scientific manuscripts and produces a comprehensive evaluation, checking theoretical results, validating experiments, suggesting improvements, and identifying potential flaws. By utilizing inference scaling techniques, PAT is able to identify deeper issues than a single model call alone, achieving a 34% improvement over zero-shot recall on mathematical errors in the SPOT benchmark. Pilot deployments of PAT as a pre-submission tool for authors at two major Computer Science conferences -- STOC and ICML -- demonstrate its ability to identify critical errors and suggest substantive improvements to research papers. By catching errors early, PAT eases the cognitive burden placed on referees, while preserving their control over the outcomes of the review process.

AI评审论文工具科学验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。