arXiv:2508.07043cs.AIcs.MA2025-08被引 11

用多智能体系统实现生物信息学分析自动化,准确率超顶尖大模型27%。

K-Dense Analyst: Towards Fully Automated Scientific Analysis

  • 分层多智能体设计,双循环架构分解任务并验证执行
  • 在BixBench上达29.2%准确率,比GPT-5高6.3个百分点
  • 适合追求自动化科学发现的科研人员与生物信息平台开发者

现代生物信息学分析的复杂性导致数据生成与科学洞察之间存在巨大鸿沟。尽管大语言模型(LLMs)在科学推理方面展现出潜力,但在需要迭代计算、工具集成和严格验证的真实分析流程中仍存在根本局限。我们提出K-Dense Analyst,一个基于分层多智能体的系统,通过双循环架构实现自主生物信息学分析。该系统作为更广泛K-Dense平台的一部分,将规划与可验证执行结合,利用专用智能体在安全计算环境中将复杂目标分解为可执行、可验证的任务。在BixBench这一开放式生物分析综合基准测试中,K-Dense Analyst达到29.2%的准确率,较表现最佳的语言模型(GPT-5)高出6.3个百分点,相较普遍认为最强的现有大模型提升近27%。值得注意的是,直接使用Gemini 2.5 Pro仅获18.3%准确率,表明我们的架构创新显著释放了远超基础模型能力的潜力。研究揭示,自主科学推理不仅需要增强型语言模型,更需量身定制的系统,以弥合高层次科学目标与底层计算执行之间的差距。这些成果标志着向全自动化计算生物学家迈出了重要一步,有望加速生命科学领域的发现进程。

原文摘要 · Abstract (English)

The complexity of modern bioinformatics analysis has created a critical gap between data generation and developing scientific insights. While large language models (LLMs) have shown promise in scientific reasoning, they remain fundamentally limited when dealing with real-world analytical workflows that demand iterative computation, tool integration and rigorous validation. We introduce K-Dense Analyst, a hierarchical multi-agent system that achieves autonomous bioinformatics analysis through a dual-loop architecture. K-Dense Analyst, part of the broader K-Dense platform, couples planning with validated execution using specialized agents to decompose complex objectives into executable, verifiable tasks within secure computational environments. On BixBench, a comprehensive benchmark for open-ended biological analysis, K-Dense Analyst achieves 29.2% accuracy, surpassing the best-performing language model (GPT-5) by 6.3 percentage points, representing nearly 27% improvement over what is widely considered the most powerful LLM available. Remarkably, K-Dense Analyst achieves this performance using Gemini 2.5 Pro, which attains only 18.3% accuracy when used directly, demonstrating that our architectural innovations unlock capabilities far beyond the underlying model's baseline performance. Our insights demonstrate that autonomous scientific reasoning requires more than enhanced language models, it demands purpose-built systems that can bridge the gap between high-level scientific objectives and low-level computational execution. These results represent a significant advance toward fully autonomous computational biologists capable of accelerating discovery across the life sciences.

自动化分析多智能体生物信息学大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。