用大模型自动标注单细胞类型,无需调优也能准确识别。
Single-Cell Omics Arena: A Benchmark Study for Large Language Models on Cell Type Annotation Using Single-Cell Data
- 用8个指令微调大模型,在11个数据集上测试细胞类型标注能力。
- 大模型在单细胞数据上分类准确,且能跨模态扩展至多组学数据。
- 链式思考提示法可生成详细生物学解释,适合生物研究者使用。
过去十年,单细胞测序技术实现了对数千个单个细胞中多种分子特征的并行分析,使科学家能够探究复杂组织的多样性功能并揭示疾病机制。其中,将细胞分配到特定类型是理解细胞异质性的基础步骤,但传统方法依赖大量人工和专业知识。近年来,大语言模型(LLMs)在处理海量文本、自动提取关键生物知识(如标记基因)方面表现突出,有望实现更高效、自动化的细胞类型注释。为全面评估现代指令微调大模型在单细胞基因组学中自动化细胞类型识别的能力,我们提出SOAR,一项涵盖8个指令微调大模型、覆盖11个数据集的综合性基准研究,涵盖多种细胞类型与物种。研究评估了大模型在单细胞RNA测序(scRNA-seq)数据中的分类与注释性能,并探索其通过跨模态翻译扩展至多组学数据的应用潜力。同时,我们检验了链式思考(CoT)提示技术在注释过程中生成详细生物学见解的有效性。结果表明,大模型无需额外微调即可提供稳健的单细胞数据分析解释,显著推进了基因组研究中细胞类型注释的自动化进程。
原文摘要 · Abstract (English)
Over the past decade, the revolution in single-cell sequencing has enabled the simultaneous molecular profiling of various modalities across thousands of individual cells, allowing scientists to investigate the diverse functions of complex tissues and uncover underlying disease mechanisms. Among all the analytical steps, assigning individual cells to specific types is fundamental for understanding cellular heterogeneity. However, this process is usually labor-intensive and requires extensive expert knowledge. Recent advances in large language models (LLMs) have demonstrated their ability to efficiently process and synthesize vast corpora of text to automatically extract essential biological knowledge, such as marker genes, potentially promoting more efficient and automated cell type annotations. To thoroughly evaluate the capability of modern instruction-tuned LLMs in automating the cell type identification process, we introduce SOAR, a comprehensive benchmarking study of LLMs for cell type annotation tasks in single-cell genomics. Specifically, we assess the performance of 8 instruction-tuned LLMs across 11 datasets, spanning multiple cell types and species. Our study explores the potential of LLMs to accurately classify and annotate cell types in single-cell RNA sequencing (scRNA-seq) data, while extending their application to multiomics data through cross-modality translation. Additionally, we evaluate the effectiveness of chain-of-thought (CoT) prompting techniques in generating detailed biological insights during the annotation process. The results demonstrate that LLMs can provide robust interpretations of single-cell data without requiring additional fine-tuning, advancing the automation of cell type annotation in genomics research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。