arXiv:2605.05985cs.AIcs.MA2026-05

构建可追溯的多智能体系统,实现精准医学研究全流程自动化

BioResearcher: Scenario-Guided Multi-Agent for Translational Medicine

  • 按场景调用版本化研究手册,分派专业子智能体协同工作
  • 在109项单步测试中通过率83.49%,临床端到端任务正向命中率达74.7%
  • 适合需要可审计、高可信医学推理的研究者和机构

精准医学需将模糊的研究目标转化为融合文献、临床试验、专利及多组学数据的证据综合,同时保留标识符、不确定性与可追溯溯源。通用基础模型和现成工具增强或多智能体系统难以胜任:它们常生成一次性答案或无止境运行,无法满足异构生物医学数据所需的可审计、场景化工作流。本文提出Ingenix BioResearcher,一个场景引导的多智能体系统,将查询映射至版本化研究手册,调度30多个工具与机器学习接口的专用子智能体,结合结构化数据库访问与沙盒代码执行进行基因组级分析,并在声明层面应用多模型一致性校验后进行编辑集成。我们在单元能力、开放性生物医学推理及端到端临床发现任务上评估该系统,在109项单步测试中通过率达83.49%(平均得分0.892),在BixBench-Verified-50上表现89.33%,在BaisBench科学发现任务中均分达0.758,端到端临床基准30个查询中正向命中率74.7%±3.3%,负向明确率96.8%±0.2%。结果表明其在单元、开放及端到端评估中均具广泛竞争力。

原文摘要 · Abstract (English)

Translational medicine turns underspecified development goals into evidence synthesis that must combine literature, trials, patents, and quantitative multi-omics analysis while preserving identifiers, uncertainty, and retrievable provenance. General-purpose foundation models and off-the-shelf tool-augmented or multi-agent systems are not built for this: they tend to produce single-shot answers or run open-endedly, and fall short on the auditable, scenario-specific workflows that heterogeneous biomedical sources demand. This paper introduces Ingenix BioResearcher, a scenario-guided multi-agent system that maps queries to versioned research playbooks, delegates to specialized subagents over 30+ tools and machine-learning endpoints, mixes structured database access with sandboxed code for genome-scale analyses, and applies claim-level multi-model reconciliation before editorial assembly. We evaluate BioResearcher across unit-level capabilities, open-ended biomedical reasoning, and end-to-end clinical discovery. It leads evaluated baselines on 109 single-step tests (83.49% pass rate; 0.892 average score), achieves strong biomedical benchmark performance (89.33% on BixBench-Verified-50 and the top 0.758 mean score on BaisBench Scientific Discovery), and leads on a 30-query clinical end-to-end benchmark with the highest positive hit rate (74.7% $\pm$ 3.3%) and negative clear rate (96.8% $\pm$ 0.2%). These results show broad, competitive performance across unit-level, open-ended, and end-to-end clinical evaluations.

精准医学多智能体生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。