arXiv:2602.11609cs.AIq-bio.GN2026-02NeurIPS被引 7

用大模型直接分析单细胞数据,自动完成细胞注释与发育轨迹推断。

scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery

  • 大模型直接读取原始单细胞数据,分步推理并修正分析过程。
  • 相比一次性提示,迭代推理使细胞注释准确率提升11%,轨迹重建误差降低30%。
  • 生成可解释的推理链条,适合生物学家和算法研究者使用。

我们提出scPilot,首个实现组学原生推理的系统:大语言模型(LLM)以自然语言对话方式,直接分析单细胞RNA测序数据及按需调用生信工具。scPilot将核心单细胞分析任务——细胞类型注释、发育轨迹重建、转录因子靶向——转化为逐步推理问题,要求模型给出解释,并在新证据下修正判断。为评估进展,我们发布scBench,包含9个专家标注的数据集与评分器,用于衡量不同LLM在组学原生推理上的表现。实验显示,o1模型通过迭代推理使细胞注释平均准确率提升11%;Gemini-2.5-Pro相比一次性提示,轨迹图编辑距离减少30%。同时生成透明推理轨迹,解释标记基因模糊性与调控逻辑。通过将大模型扎根于原始组学数据,scPilot实现可审计、可解释、具诊断价值的单细胞分析。代码、数据与包已开源。

原文摘要 · Abstract (English)

We present scPilot, the first systematic framework to practice omics-native reasoning: a large language model (LLM) converses in natural language while directly inspecting single-cell RNA-seq data and on-demand bioinformatics tools. scPilot converts core single-cell analyses, i.e., cell-type annotation, developmental-trajectory reconstruction, and transcription-factor targeting, into step-by-step reasoning problems that the model must solve, justify, and, when needed, revise with new evidence. To measure progress, we release scBench, a suite of 9 expertly curated datasets and graders that faithfully evaluate the omics-native reasoning capability of scPilot w.r.t various LLMs. Experiments with o1 show that iterative omics-native reasoning lifts average accuracy by 11% for cell-type annotation and Gemini-2.5-Pro cuts trajectory graph-edit distance by 30% versus one-shot prompting, while generating transparent reasoning traces explain marker gene ambiguity and regulatory logic. By grounding LLMs in raw omics data, scPilot enables auditable, interpretable, and diagnostically informative single-cell analyses. Code, data, and package are available at https://github.com/maitrix-org/scPilot

单细胞分析大模型推理组学原生可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。