用大模型解析单细胞基因数据,跨物种分析更准更快
scReader: Prompting Large Language Models to Interpret scRNA-seq Data
- 将基因功能描述与大模型结合,生成细胞级表达表征
- 在人鼠发育细胞上实现高精度注释,准确率显著提升
- 适合生物信息学研究者做跨物种单细胞数据分析
大型语言模型(LLMs)在建模文本序列隐含关系方面表现卓越,为生命科学领域带来新机遇。尽管多物种单细胞组学数据量庞大,但不同物种间数据规模差异大,制约了跨物种遗传数据解析模型的发展。本文提出一种混合方法,融合大模型的通用知识能力与单细胞组学的领域特异性表征模型。以基因作为基本表征单元,利用成熟语言模型(如LLaMA-2)的功能描述初始化基因表征。通过输入单细胞基因表达数据并设计提示词,基于不同物种和细胞类型中基因的差异表达水平,有效构建细胞表征。实验中构建了人和小鼠的发育细胞数据,重点针对难标注的细胞类型。评估任务包括细胞注释与可视化分析。结果表明,该方法相比现有基于大模型的方法在准确性和跨物种兼容性上均有显著提升。该混合策略增强了单细胞数据表征能力,为未来跨物种遗传分析提供了稳健框架。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable advancements, primarily due to their capabilities in modeling the hidden relationships within text sequences. This innovation presents a unique opportunity in the field of life sciences, where vast collections of single-cell omics data from multiple species provide a foundation for training foundational models. However, the challenge lies in the disparity of data scales across different species, hindering the development of a comprehensive model for interpreting genetic data across diverse organisms. In this study, we propose an innovative hybrid approach that integrates the general knowledge capabilities of LLMs with domain-specific representation models for single-cell omics data interpretation. We begin by focusing on genes as the fundamental unit of representation. Gene representations are initialized using functional descriptions, leveraging the strengths of mature language models such as LLaMA-2. By inputting single-cell gene-level expression data with prompts, we effectively model cellular representations based on the differential expression levels of genes across various species and cell types. In the experiments, we constructed developmental cells from humans and mice, specifically targeting cells that are challenging to annotate. We evaluated our methodology through basic tasks such as cell annotation and visualization analysis. The results demonstrate the efficacy of our approach compared to other methods using LLMs, highlighting significant improvements in accuracy and interoperability. Our hybrid approach enhances the representation of single-cell data and offers a robust framework for future research in cross-species genetic analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。