arXiv:2508.13201q-bio.GNcs.AI2025-08被引 6

首个针对单细胞组学分析的AI智能体综合评测系统。

Benchmarking LLM-based agents for single-cell omics analysis

  • 构建统一平台与50个真实任务,评估智能体多维度能力。
  • 多智能体协作显著提升效率,代码生成是成功关键。
  • 适合计算生物学、AI医疗交叉研究者参考。

单细胞组学数据激增暴露了传统手动分析流程的局限性。AI智能体通过自适应规划、可执行代码生成、可追溯决策和实时知识融合带来范式变革,但缺乏全面评测体系严重制约进展。本文提出全新评测系统,包含兼容多种智能体框架和大模型的统一平台、涵盖认知程序合成、协作能力、执行效率、生物信息学知识融合及任务完成质量的多维指标,以及覆盖多组学、物种和测序技术的50个真实世界分析任务。评估显示,Grok3-beta在测试框架中表现最优;多智能体架构通过角色分工显著优于单智能体方案。能力归因分析表明,高质量代码生成对任务成功至关重要,自我反思影响最大,其次为检索增强生成(RAG)与规划能力。研究揭示代码生成、长上下文处理和上下文感知知识检索仍是主要挑战,为开发稳健的计算生物学AI智能体提供了关键实证基础与最佳实践。

原文摘要 · Abstract (English)

Background: The surge in single-cell omics data exposes limitations in traditional, manually defined analysis workflows. AI agents offer a paradigm shift, enabling adaptive planning, executable code generation, traceable decisions, and real-time knowledge fusion. However, the lack of a comprehensive benchmark critically hinders progress. Results: We introduce a novel benchmarking evaluation system to rigorously assess agent capabilities in single-cell omics analysis. This system comprises: a unified platform compatible with diverse agent frameworks and LLMs; multidimensional metrics assessing cognitive program synthesis, collaboration, execution efficiency, bioinformatics knowledge integration, and task completion quality; and 50 diverse real-world single-cell omics analysis tasks spanning multi-omics, species, and sequencing technologies. Our evaluation reveals that Grok3-beta achieves state-of-the-art performance among tested agent frameworks. Multi-agent frameworks significantly enhance collaboration and execution efficiency over single-agent approaches through specialized role division. Attribution analyses of agent capabilities identify that high-quality code generation is crucial for task success, and self-reflection has the most significant overall impact, followed by retrieval-augmented generation (RAG) and planning. Conclusions: This work highlights persistent challenges in code generation, long-context handling, and context-aware knowledge retrieval, providing a critical empirical foundation and best practices for developing robust AI agents in computational biology.

AI智能体单细胞组学评测基准计算生物学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。