用大模型实现跨组织、发现新细胞类型的通用注释框架
scAgent: Universal Single-Cell Annotation via a LLM Agent
- 基于大语言模型构建智能代理,自动识别和发现细胞类型
- 在160种细胞类型、35个组织上表现优于现有方法
- 数据效率高,能快速学习新细胞类型,适合生物医学研究者
细胞类型注释对理解细胞异质性至关重要。基于单细胞RNA测序数据和深度学习模型,已有研究在特定组织内固定数量的细胞类型注释上取得进展。然而,能够跨组织泛化、发现新细胞类型并扩展至未知细胞类型的通用细胞注释仍缺乏探索。为此,本文提出scAgent——一种基于大语言模型(LLM)的通用细胞注释框架。scAgent可在多种组织中识别细胞类型并发现新细胞类型;此外,其具有高效的数据利用能力,能快速学习新细胞类型。在160种细胞类型和35个组织上的实验表明,scAgent在通用细胞类型注释、新细胞发现及对新细胞类型的可扩展性方面均表现优异。
原文摘要 · Abstract (English)
Cell type annotation is critical for understanding cellular heterogeneity. Based on single-cell RNA-seq data and deep learning models, good progress has been made in annotating a fixed number of cell types within a specific tissue. However, universal cell annotation, which can generalize across tissues, discover novel cell types, and extend to novel cell types, remains less explored. To fill this gap, this paper proposes scAgent, a universal cell annotation framework based on Large Language Models (LLMs). scAgent can identify cell types and discover novel cell types in diverse tissues; furthermore, it is data efficient to learn novel cell types. Experimental studies in 160 cell types and 35 tissues demonstrate the superior performance of scAgent in general cell-type annotation, novel cell discovery, and extensibility to novel cell type.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。