用30个语言模型代理自动完成高阶宇宙学研究任务
Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery
- 构建30个专精分工的语言代理,实现科研全流程自动化
- 在超新星数据中成功测量宇宙学参数,性能优于现有顶尖模型
- 完全无需人工干预,代码可本地运行,开源且支持云端部署
我们提出一个用于自动化科学探索任务的多智能体系统 cmbagent(https://github.com/CMBAgents/cmbagent),由约30个大语言模型代理组成,采用规划与控制策略协调智能体工作流,全程无需人工介入。每个代理专注于不同任务(如文献与代码库检索、代码编写、结果解读、输出批判),系统可本地执行代码。我们成功将 cmbagent 应用于博士级宇宙学任务(利用超新星数据测量宇宙学参数),并在两个基准测试集上评估其表现,结果优于当前最先进大模型。源代码已公开于 GitHub,演示视频亦可获取,系统已部署于 HuggingFace 并将上线云服务。
原文摘要 · Abstract (English)
We present a multi-agent system for automation of scientific research tasks, cmbagent (https://github.com/CMBAgents/cmbagent). The system is formed by about 30 Large Language Model (LLM) agents and implements a Planning & Control strategy to orchestrate the agentic workflow, with no human-in-the-loop at any point. Each agent specializes in a different task (performing retrieval on scientific papers and codebases, writing code, interpreting results, critiquing the output of other agents) and the system is able to execute code locally. We successfully apply cmbagent to carry out a PhD level cosmology task (the measurement of cosmological parameters using supernova data) and evaluate its performance on two benchmark sets, finding superior performance over state-of-the-art LLMs. The source code is available on GitHub, demonstration videos are also available, and the system is deployed on HuggingFace and will be available on the cloud.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。