用基因描述文本生成细胞嵌入,揭示运动神经元脆弱性机制
Bridging Large Language Models and Single-Cell Transcriptomics in Dissecting Selective Motor Neuron Vulnerability
- 用高表达基因的文本注释生成语义丰富的细胞嵌入
- 在多个数据集中实现更精准的细胞类型聚类与脆弱性分析
- 适合生物信息学与神经退行性疾病研究者使用
单细胞转录组数据中理解细胞身份与功能仍是计算生物学的关键挑战。本文提出一种新框架,利用NCBI Gene数据库中的基因特异性文本注释,生成生物学上下文相关的细胞嵌入。对每个单细胞RNA测序(scRNA-seq)数据集中的细胞,按基因表达水平排序,提取前N个高表达基因的NCBI Gene描述,并通过大语言模型(LLMs)将其转化为向量嵌入。使用的模型包括OpenAI text-embedding-ada-002、text-embedding-3-small、text-embedding-3-large(2024年1月版),以及领域专用模型BioBERT和SciBERT。嵌入通过各基因表达加权平均计算,获得紧凑且语义丰富的表示。该多模态策略融合结构化生物数据与前沿语言模型,提升了下游任务的可解释性,如细胞类型聚类、细胞脆弱性解析和轨迹推断。
原文摘要 · Abstract (English)
Understanding cell identity and function through single-cell level sequencing data remains a key challenge in computational biology. We present a novel framework that leverages gene-specific textual annotations from the NCBI Gene database to generate biologically contextualized cell embeddings. For each cell in a single-cell RNA sequencing (scRNA-seq) dataset, we rank genes by expression level, retrieve their NCBI Gene descriptions, and transform these descriptions into vector embedding representations using large language models (LLMs). The models used include OpenAI text-embedding-ada-002, text-embedding-3-small, and text-embedding-3-large (Jan 2024), as well as domain-specific models BioBERT and SciBERT. Embeddings are computed via an expression-weighted average across the top N most highly expressed genes in each cell, providing a compact, semantically rich representation. This multimodal strategy bridges structured biological data with state-of-the-art language modeling, enabling more interpretable downstream applications such as cell-type clustering, cell vulnerability dissection, and trajectory inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。