AI团队自主协作,长期探索科学实验,避免重复试错。
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

- 多个AI智能体自主组队,围绕候选假设协同工作。
- 在生物医学等任务中平均得分提升至74.4%,优于最强基线8.33%。
- 适合需要长期迭代、多方向探索的复杂科研场景。
科学研究依赖于假设生成、实验设计、执行与修正的迭代过程。现有AI方法通常沿单一研究路径运行或通过中心化规划者协调,目标固定,难以持续并行探索、随证据变化调整,也无法长期保留失败路径的知识。本文提出AutoScientists,一种用于长期计算型科学实验的去中心化AI智能体团队。智能体共享实验状态,自主组织为围绕高潜力假设的小组,对提案进行批判性评估后再使用计算资源,并共享成功与失败以减少重复探索。在相同实验预算下,AutoScientists在生物医学机器学习、语言模型训练优化和蛋白质适应度预测任务中均超越先前AI方法。在涵盖生物医学影像、蛋白工程、单细胞组学和药物发现的BioML-Bench上,24项任务平均排行榜百分位达74.4%,比最强基线提升8.33%。在GPT训练优化任务中,其达到目标验证比特/字的效率是Autoresearch的1.9倍,且在初始优胜模型基础上继续发现7项改进,而单智能体方法未发现任何有效改进(0项)。在ProteinGym蛋白质适应度预测中,该方法使ACE2-刺突结合预测的斯皮尔曼相关系数提升12.5%,跨全部217个测试集平均提升6.5%。
原文摘要 · Abstract (English)
Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existing approaches typically follow a single research trajectory or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration, adapt as experimental evidence changes, or preserve knowledge of failed directions over long-running experiments. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Agents interpret a shared experimental state, self-organize into teams around promising hypotheses, critique proposals before using experimental compute, and share successes and failures to reduce redundant exploration. Under matched experimental budgets, AutoScientists improves over prior AI agents across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest AI agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than Autoresearch and continues discovering improvements from a starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-Spike binding that improves over the current state-of-the-art model by +12.5% in Spearman correlation. Applied without modification across all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% (Spearman correlation).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。