AI科学家新范式:多智能体系统实现分钟级交互式科研探索
Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery
- 构建多智能体系统,通过持续世界状态实现跨轮次研究上下文保持
- 在BixBench生物计算基准上达到64.4%准确率,超越基线14-26个百分点
- 支持人机协同与全自动化两种模式,适合需要快速迭代的科研团队
面向科学发现的人工智能系统展现出巨大潜力,但现有方法多为专有系统且采用批量处理模式,每轮研究需数小时,难以实现实时研究人员干预。本文提出Deep Research,一种多智能体系统,支持以分钟为单位完成的交互式科学研究。该架构包含规划、数据分析、文献检索和新颖性检测等专用智能体,通过持久化世界状态统一各智能体,保持跨迭代研究上下文。系统提供两种运行模式:半自动模式支持选择性人工检查点,全自动化模式用于长期调查。在BixBench计算生物学基准上的评估显示,其在开放式回答任务中达到48.8%准确率,在多项选择任务中达64.4%,相比现有基线提升14至26个百分点。对开放文献获取限制及自动化新颖性评估固有挑战的分析,为人工智能辅助科研流程的实际部署提供了参考。
原文摘要 · Abstract (English)
Artificial intelligence systems for scientific discovery have demonstrated remarkable potential, yet existing approaches remain largely proprietary and operate in batch-processing modes requiring hours per research cycle, precluding real-time researcher guidance. This paper introduces Deep Research, a multi-agent system enabling interactive scientific investigation with turnaround times measured in minutes. The architecture comprises specialized agents for planning, data analysis, literature search, and novelty detection, unified through a persistent world state that maintains context across iterative research cycles. Two operational modes support different workflows: semi-autonomous mode with selective human checkpoints, and fully autonomous mode for extended investigations. Evaluation on the BixBench computational biology benchmark demonstrated state-of-the-art performance, achieving 48.8% accuracy on open response and 64.4% on multiple-choice evaluation, exceeding existing baselines by 14 to 26 percentage points. Analysis of architectural constraints, including open access literature limitations and challenges inherent to automated novelty assessment, informs practical deployment considerations for AI-assisted scientific workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。