AI可自主完成从构思到发表的完整科研流程,通过自动生成论文并过审。
Towards End-to-End Automation of AI Research

- 构建智能体系统,全程自主生成研究想法、代码与论文。
- 在机器学习会议研讨会上,生成论文通过首轮同行评审,录用率达70%。
- 支持有模板和无模板两种模式,适用于定向与开放探索研究。
科学自动化是人工智能领域长期追求的目标。尽管社区已在科学流程的个别环节取得显著进展,但能够端到端自主完成从研究构想到论文发表全过程的系统仍遥不可及。本文提出迄今最强有力的示范:人工智能科学家(The AI Scientist),可自主生成研究构想、编写代码、执行实验、绘制分析数据、撰写完整科学论文,并进行自我同行评审。其生成的研究成果质量足以使由AI系统生成的论文通过主流机器学习会议研讨会的第一轮同行评审,该研讨会接受率为70%。该系统利用现代基础模型构建复杂智能体架构,在两种场景下验证:一种是基于人类提供的代码模板的聚焦模式,另一种是无模板、开放式探索的智能体搜索模式。两种模式均能生成多样研究想法,并自动测试、报告与评估。这一成果彰显了人工智能在科学研究中日益增强的贡献能力,预示着研究范式的潜在变革。然而,此类技术也带来挑战,如可能加重审稿负担、增加文献噪声。若负责任地发展,此类自主系统或可极大加速科学发现。
原文摘要 · Abstract (English)
The automation of science is a long-standing ambition in the field of AI. While the community has made significant progress in automating individual components of the scientific process, a system that autonomously navigates the entire research lifecycle -- from conception to publication -- has remained out of reach. Here, we present the strongest demonstration to date toward automating the entire process end-to-end. We present The AI Scientist, which creates research ideas, writes code, runs experiments, plots and analyzes data, writes the entire scientific manuscript and performs its own peer review. Its ideas, execution, and presentation are of sufficient quality to produce a manuscript generated by an AI system that passes the first round of peer review at a major machine learning conference workshop. The workshop has an acceptance rate of 70 percent. Our system leverages modern foundation models within a complex agentic system. We evaluate The AI Scientist in two settings: a focused mode using human-provided code templates as an initial scaffold to conduct research on a specific topic, and a template-free, open-ended mode that leverages agentic search for wider scientific exploration. Both settings produce diverse ideas and automatically test, report on, and evaluate them. This achievement demonstrates AI's growing capacity for scientific contribution and signifies a potential paradigm shift in how research is conducted. As with any impactful new technology, there could be significant risks, including taxing overwhelmed review systems and adding noise to scientific literature. However, if developed responsibly, such autonomous systems could greatly accelerate scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。