用可追溯的AI代理系统,自动完成脑科学跨尺度研究任务。
BrainPilot: Automating Brain Discovery with Agentic Research

- 多智能体协作,基于7233项知识条目和72个方法模块开展研究。
- 全程记录工作流,关键步骤可审计,错误率低于10%。
- 适合需要高可靠性的脑科学研究员和自动化研究团队。
理解大脑日益依赖于跨尺度、多模态和跨学科证据的整合。解决单一科研问题需协调一系列操作,从文献调研到分析执行及结合领域知识解读结果。当前AI代理缺乏脑科学领域专长,可能虚构结论、推理漂移,且缺乏专家介入点。这些缺陷在脑科学中尤为严重,因结论直接影响后续科学判断,依赖实验室特定知识与严谨的人类判断。本文提出完全开源的多智能体系统BrainPilot,通过可追溯日志和代理验证结果加速脑科学研究。主研人(PI)代理协调基于精选领域知识的专家代理:统一脑科学知识库包含7,233项索引条目,技能库含7个研究领域中的72个可复用方法单元。每个关键步骤均记录于“追踪图”中,可审计地关联子目标、工具使用、证据与主张,支持研究者追踪与审查流程。审计代理进一步集成防虚构检测。评估中,我们在三个来自Agents' Last Exam的任务上测试,并引入自研基准BrainPilotBench-v0及端到端案例研究。结果表明,采用开源骨干模型的BrainPilot性能媲美顶尖代理框架,成本更低。
原文摘要 · Abstract (English)
Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined points for expert intervention. These failures are especially costly in brain science, where conclusions feed into downstream scientific claims and depend on laboratory-specific expertise and careful human judgment. We present \textbf{BrainPilot} a \textbf{fully open-source} multi-agent system that accelerates brain science research with traceable logs and agent-verified results. A principal investigator (PI) agent coordinates specialist agents grounded in curated domain knowledge: a unified brain science knowledge base containing 7{,}233 indexed items and a skill library of 72 reusable methodology units across seven research domains. Every major step is recorded in the Graph of Trace, an auditable record that links subgoals, tool use, evidence, and claims and allows researchers to follow and inspect the workflow. An Auditor agent further integrates fabrication checking into the workflow. For evaluation, we run three brain science tasks from Agents' Last Exam, introduce our own benchmark, \textbf{BrainPilotBench-v0}, and present additional end-to-end case studies. Across these evaluations, BrainPilot with an open-source backbone model attains performance comparable to state-of-the-art agent framework with less costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。