arXiv:2606.13007cs.LGcs.AI2026-06

用基因知识增强的多模态模型,提升单细胞测序聚类准确性

scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing

论文配图:scLLM-DSC: LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering for Single-Cell RNA Sequencing
图 1 · 摘自论文原文
  • 融合基因功能知识与表达数据结构,构建语义-拓扑双视角表征
  • 在11个基准上显著优于现有方法,准确率大幅提升
  • 适合需要精准细胞类型识别的研究者使用

单细胞RNA测序分析中,聚类是识别细胞群体、解析组织异质性的关键步骤。现有方法主要依赖数值统计模式,忽视基因编码的内在生物学功能。虽然大语言模型具备良好语义能力,但其生成式预训练目标与聚类等判别任务存在结构不匹配问题。为此,我们提出scLLM-DSC:一种基于大语言模型知识的跨模态深度结构聚类框架。该方法通过两个视角构建语义基础表征:一是基于NCBI基因先验和上下文化的Cell2Sentence嵌入的知识驱动语义视图;二是通过图引导编码器提取的结构感知拓扑视图。关键在于引入跨模态对比对齐机制,在统一潜在空间中强制生物语义与转录组特征的一致性。大量实验表明,scLLM-DSC在11个主流基准上显著超越现有最优方法,聚类准确率明显提升。

原文摘要 · Abstract (English)

Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity. However, existing methods focus on mining numerical statistical patterns, suffering from semantic agnosticism by neglecting the intrinsic biological functions encoded by genes. While Large Language Models (LLMs) offer promising semantic capabilities, their direct adaptation to cell clustering is hindered by the structural mismatch between generative pre-training objectives and discriminative downstream tasks. To bridge this gap, we propose scLLM-DSC, a novel LLM-Knowledge Enhanced Cross-Modal Deep Structural Clustering framework. Diverging from data-driven paradigms, scLLM-DSC establishes a semantically-grounded representation by synergizing two views: a Knowledge-Driven Semantic View derived from NCBI gene priors and contextualized Cell2Sentence embeddings, and a Structure-Aware Topological View extracted via a graph-guided encoder. Crucially, we introduce a cross-modal contrastive alignment mechanism to enforce consistency between biological semantics and transcriptomic features within a unified latent space. Extensive benchmarks demonstrate that scLLM-DSC significantly outperforms eleven state-of-the-art baselines in clustering accuracy.

单细胞聚类大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。