用AI自主发现科学规律,通过可证伪机制验证新发现。
AIGS: Generating Science from AI-Powered Automated Falsification
- 设计多智能体系统,让AI全程自主完成科研流程。
- 引入可证伪机制,使AI能主动检验并验证科学假设。
- 在3个任务中初步实现有意义的科学发现,适合研究AI科研自动化者。
人工智能的快速发展显著加速了科学发现进程。基于大规模观测数据训练的深度神经网络以端到端方式提取潜在规律,辅助人类在未见场景中进行高精度预测。近年来,大语言模型(LLMs)和自主代理的兴起,使科学家能在文献综述、研究构思、实验实现和论文撰写等阶段获得交互式帮助。然而,基于基础模型的全链路自主AI研究人员仍处于起步阶段。本文研究 extbf{AI生成科学}(AIGS),即智能体独立且自主完成整个研究过程并发现科学定律。我们重新审视科学研究的定义,认为 extit{证伪}是人类研究过程与AIGS系统设计的核心。从证伪视角看,已有系统或缺少关键环节,或严重依赖现有验证引擎,限制了其在特定领域外的应用。本文提出Baby-AIGS,作为全链路AIGS系统的初步演示,是一个多智能体系统,各智能体代表研究流程中的关键角色。通过引入FalsificationAgent,识别并验证可能的科学发现,系统具备显式的证伪能力。三个任务的初步实验表明,Baby-AIGS能够产生有意义的科学发现,尽管尚未达到资深人类研究者的水平。最后,本文深入讨论当前Baby-AIGS的局限性、可操作建议及相关伦理问题。
原文摘要 · Abstract (English)
Rapid development of artificial intelligence has drastically accelerated the development of scientific discovery. Trained with large-scale observation data, deep neural networks extract the underlying patterns in an end-to-end manner and assist human researchers with highly-precised predictions in unseen scenarios. The recent rise of Large Language Models (LLMs) and the empowered autonomous agents enable scientists to gain help through interaction in different stages of their research, including but not limited to literature review, research ideation, idea implementation, and academic writing. However, AI researchers instantiated by foundation model empowered agents with full-process autonomy are still in their infancy. In this paper, we study $\textbf{AI-Generated Science}$ (AIGS), where agents independently and autonomously complete the entire research process and discover scientific laws. By revisiting the definition of scientific research, we argue that $\textit{falsification}$ is the essence of both human research process and the design of an AIGS system. Through the lens of falsification, prior systems attempting towards AI-Generated Science either lack the part in their design, or rely heavily on existing verification engines that narrow the use in specialized domains. In this work, we propose Baby-AIGS as a baby-step demonstration of a full-process AIGS system, which is a multi-agent system with agents in roles representing key research process. By introducing FalsificationAgent, which identify and then verify possible scientific discoveries, we empower the system with explicit falsification. Experiments on three tasks preliminarily show that Baby-AIGS could produce meaningful scientific discoveries, though not on par with experienced human researchers. Finally, we discuss on the limitations of current Baby-AIGS, actionable insights, and related ethical issues in detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。