arXiv:2607.28631cs.AIcs.CL2026-07综述被引 1

用AI自动评审AI科研成果,首次量化对比多个自研系统表现。

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

论文配图:Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
图 1 · 摘自论文原文
  • 用多个顶级大模型自动评审论文,从原创性、严谨性等四方面打分。
  • 商业AI科研系统FARS得分远超其他框架,是次优系统的2倍以上。
  • 多模型评审结果一致,证明自动化评估可靠且可扩展。

具备自主研究能力的AI Scientist系统有望显著加速科学发现,但评估和比较其生成论文的质量仍是开放挑战。本文提出并实施一套严格的基准测试协议,利用前沿大语言模型组成的自动化同行评审系统,从原创性、科学严谨性、清晰度和重要性四个核心维度评估论文。我们对四个领先AI Scientist框架(Sakana AI v1/v2、CycleResearcher、Data-to-Paper)进行测试,每个在15个由商业公司FARS发布的研究提案上生成论文,共得60篇,与15篇FARS基准论文一同评估。使用GPT-5.4、Gemini和Claude三名独立大模型评审员发现,FARS基准论文平均得分2.14–2.47(1–5分制),显著高于其他系统(1.00–1.87)。其中,FARS在Gemini和Claude评分中均超过次优系统2倍以上。Gemini与Claude评分高度一致(ρ=0.907, p<0.001),且与综合评分强相关(ρ=0.961, p<0.001),验证了自动化评估的可靠性。但GPT-5.4一致性较低(ρ≈0.32),表明其评价标准不同。该研究建立了首个针对AI Scientist系统的量化基准,证明多模型评审是可扩展、一致的自主研究质量评估框架。

原文摘要 · Abstract (English)

AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generated papers remains an open challenge. We propose and implement a rigorous benchmarking protocol using an automated peer-review system that harnesses frontier large language models to assess scientific papers across four core dimensions: originality, scientific rigor, clarity, and significance. We evaluate four leading AI Scientist frameworks: \textit{Sakana AI (v1 & v2)}, \textit{CycleResearcher}, and \textit{Data-to-Paper}. Each framework was run on a consistent set of 15 research proposals published by a commercial autonomous AI scientist company (FARS), generating 60 papers that we evaluate alongside 15 FARS benchmark papers. Using three independent LLM reviewers (GPT-5.4, Gemini, and Claude), we find that FARS benchmark papers significantly outperform all competing frameworks, achieving mean scores of 2.14--2.47 on a 1--5 scale compared to 1.00--1.87 for other systems. Notably, FARS scores are more than 2$\times$ higher than the next-best systems on Gemini and Claude evaluations. We find strong agreement among Gemini and Claude ($ρ$ = 0.907, $p < 0.001$), and both correlate extremely strongly with the synthesis score ($ρ$ = 0.961, $p < 0.001$), validating the reliability of automated evaluation. However, GPT-5.4 exhibits weaker agreement ($ρ\approx 0.32$), suggesting it evaluates papers using different criteria. These results establish the first quantitative benchmark for AI Scientist systems and demonstrate that multi-model LLM evaluation provides a scalable, consistent framework for assessing autonomous research quality.

AI科研自动化评估大模型评审基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。