AI系统自主探索科学难题,逐步超越人类顶尖成果。
DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- 将科学发现建模为贝叶斯优化,通过假设-验证-分析循环推进
- 生成约5000个新想法,验证1100个,三项任务性能超人类SOTA
- 适合对自主科研、智能实验设计感兴趣的学者与开发者
现有AI科学家系统虽能生成新发现,但常缺乏解决紧迫人类挑战的聚焦性。本文提出DeepScientist,一个可自主运行长达数月的定向科学发现系统。它将发现过程形式化为贝叶斯优化问题,通过层级评估机制实现‘假设-验证-分析’的闭环。利用累积的发现记忆,系统智能平衡探索与利用,选择最具潜力的发现推进至更高精度验证。该系统消耗超过20,000 GPU小时,生成约5,000个独特科学构想,实验验证约1,100个,最终在三个前沿人工智能任务上分别超越人类设计的最先进方法183.7%、1.9%和7.9%。本工作首次提供大规模证据,证明AI可在科学任务中持续超越人类最先进水平,产出真正推动科学边界的实质性成果。为促进后续研究,所有实验日志与系统代码将在https://github.com/ResearAI/DeepScientist/ 开源。
原文摘要 · Abstract (English)
While previous AI Scientist systems can generate novel findings, they often lack the focus to produce scientifically valuable contributions that address pressing human-defined challenges. We introduce DeepScientist, a system designed to overcome this by conducting goal-oriented, fully autonomous scientific discovery over month-long timelines. It formalizes discovery as a Bayesian Optimization problem, operationalized through a hierarchical evaluation process consisting of "hypothesize, verify, and analyze". Leveraging a cumulative Findings Memory, this loop intelligently balances the exploration of novel hypotheses with exploitation, selectively promoting the most promising findings to higher-fidelity levels of validation. Consuming over 20,000 GPU hours, the system generated about 5,000 unique scientific ideas and experimentally validated approximately 1100 of them, ultimately surpassing human-designed state-of-the-art (SOTA) methods on three frontier AI tasks by 183.7\%, 1.9\%, and 7.9\%. This work provides the first large-scale evidence of an AI achieving discoveries that progressively surpass human SOTA on scientific tasks, producing valuable findings that genuinely push the frontier of scientific discovery. To facilitate further research into this process, we will open-source all experimental logs and system code at https://github.com/ResearAI/DeepScientist/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。