通过动态图掩码学习,提升单细胞测序聚类准确性
scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering

- 用基因掩码扰动细胞身份,构建动态图结构
- 在12个数据集上平均NMI和ARI均超越10个基线方法
- 适合处理高维稀疏、零膨胀的单细胞数据
单细胞RNA测序聚类对识别细胞类型至关重要,但高维度、稀疏性、基因缺失和噪声干扰表达表征与细胞图构建。现有掩码自编码器多依赖表达重建,图聚类方法则通常使用固定KNN图且不将恢复的表达反馈至图优化。我们提出scKDGM,一种KAN引导的动态图掩码学习框架用于scRNA-seq聚类。scKDGM采用图感知分布保持基因掩码(GDP-Mask)扰动细胞身份,基于KAN的TAKGCN编码器学习掩码视图表征,通过掩码引导表达恢复构建动态图,并利用跨视图对比学习将恢复信号传递至拓扑更新。采用ZINB损失建模过度离散与零膨胀。在12个真实scRNA-seq数据集上的实验表明,scKDGM在平均NMI和ARI上均优于10个基线方法。
原文摘要 · Abstract (English)
Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression representation and cell graph construction. Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do not feed recovered expression back into graph optimization. We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering. scKDGM uses graph-aware distribution preserving gene masking (GDP-Mask) to perturb cell identity, a KAN-based TAKGCN encoder to learn masked-view representations, mask-guided expression recovery to construct a dynamic graph, and cross-view contrastive learning to transfer recovery signals into topology updates. A ZINB loss models overdispersion and zero inflation. Experiments on 12 real scRNA-seq datasets show that scKDGM outperforms 10 baselines in average NMI and ARI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。