arXiv:2502.14297cs.IRcs.AI2025-02被引 34

评测AI科学家:自动化科研有突破但漏洞多,速度惊人却质量堪忧。

Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?

  • 用自动文献综述与实验生成实现研究流程自动化
  • 42%实验因代码错误失败,论文平均仅5篇引用且多数过时
  • 虽像本科生草稿,但3.5小时产出一文成本仅6-15美元

迈向通用人工智能(AGI)与超智能的关键一步是让AI具备自主科研能力,即所谓人工研究智能(ARI)。若机器能自主提出假设、执行实验并撰写论文,将彻底改变科学范式。Sakana近期推出‘AI科学家’,宣称可实现完全自主研究,引发广泛关注。然而,独立评估揭示其严重缺陷:文献综述对创新性判断差,常将已有概念误判为新颖;实验执行中42%因代码错误失败,结果常不实或误导;代码修改仅平均增加8%字符,适应性弱;生成论文平均仅5篇引用,34篇中仅5篇为2020年后文献,结构问题频发,如缺图、重复段落、占位符‘Conclusions Here’,部分含虚构数据。尽管如此,该系统仍代表研究自动化重大进展,能在极低人力投入下(3.5小时)以6-15美元成本生成完整论文,虽质量似仓促本科论文,但速度与效率远超传统研究者。

原文摘要 · Abstract (English)

A major step toward Artificial General Intelligence (AGI) and Super Intelligence is AI's ability to autonomously conduct research - what we term Artificial Research Intelligence (ARI). If machines could generate hypotheses, conduct experiments, and write research papers without human intervention, it would transform science. Sakana recently introduced the 'AI Scientist', claiming to conduct research autonomously, i.e. they imply to have achieved what we term Artificial Research Intelligence (ARI). The AI Scientist gained much attention, but a thorough independent evaluation has yet to be conducted. Our evaluation of the AI Scientist reveals critical shortcomings. The system's literature reviews produced poor novelty assessments, often misclassifying established concepts (e.g., micro-batching for stochastic gradient descent) as novel. It also struggles with experiment execution: 42% of experiments failed due to coding errors, while others produced flawed or misleading results. Code modifications were minimal, averaging 8% more characters per iteration, suggesting limited adaptability. Generated manuscripts were poorly substantiated, with a median of five citations, most outdated (only five of 34 from 2020 or later). Structural errors were frequent, including missing figures, repeated sections, and placeholder text like 'Conclusions Here'. Some papers contained hallucinated numerical results. Despite these flaws, the AI Scientist represents a leap forward in research automation. It generates full research manuscripts with minimal human input, challenging expectations of AI-driven science. Many reviewers might struggle to distinguish its work from human researchers. While its quality resembles a rushed undergraduate paper, its speed and cost efficiency are unprecedented, producing a full paper for USD 6 to 15 with 3.5 hours of human involvement, far outpacing traditional researchers.

AI科研自动化研究生成论文AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。