用AI代理自动做严谨实验,准确率提升3.4倍
Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents
- 设计三模块框架,分别保障可靠性、控制性和可解释性
- 在46个跨领域问题上,正确率是强基线的3.4倍
- 适合想自动化科研实验的研究者和工程团队
科学实验是人类进步的核心,需具备可靠性、系统性控制和可解释性才能得出有意义结论。尽管大语言模型在自动化科研流程方面能力增强,但实现严谨实验自动化仍具挑战。为此,我们提出Curie——一个将严谨性嵌入实验过程的AI代理框架,包含三个关键组件:内部严谨性模块以提升可靠性,跨代理严谨性模块以维持系统控制,以及实验知识模块以增强可解释性。为评估Curie,我们构建了一个新基准,涵盖四个计算机科学领域共46个问题,源自有影响力的论文和广泛使用的开源项目。与最强基线相比,我们在正确回答实验问题上实现了3.4倍的提升。Curie已在https://github.com/Just-Curieous/Curie 开源。
原文摘要 · Abstract (English)
Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap, we propose Curie, an AI agent framework designed to embed rigor into the experimentation process through three key components: an intra-agent rigor module to enhance reliability, an inter-agent rigor module to maintain methodical control, and an experiment knowledge module to enhance interpretability. To evaluate Curie, we design a novel experimental benchmark composed of 46 questions across four computer science domains, derived from influential research papers, and widely adopted open-source projects. Compared to the strongest baseline tested, we achieve a 3.4$\times$ improvement in correctly answering experimental questions. Curie is open-sourced at https://github.com/Just-Curieous/Curie.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。