通过融合主成分与判别分析,提升单细胞测序的细胞类型分类准确率。
Lower-dimensional projections of cellular expression improves cell type classification from single-cell RNA sequencing
- 融合PCA与多判别分析生成低维投影,兼顾数据方差与类别区分度。
- 在4个数据集上达98.91%准确率,未知细胞类型预测准确率达99.52%。
- 方法简单高效,不增加计算开销,适合科研与临床应用。
单细胞RNA测序(scRNA-seq)使研究人员能够在单细胞水平上研究细胞多样性,为发育过程和人类器官发生等生物机制提供全局视角。已有多种基于统计、机器学习和深度学习的方法用于细胞类型分类。大多数方法依赖于从大规模参考数据中获得的无监督低维投影。本文提出一种基于参考数据的细胞类型分类方法EnProCell:首先,通过主成分分析(PCA)与多判别分析(MDA)的集成,计算同时捕捉高方差和类间可分性的低维投影;其次,在低维表示上训练深度神经网络进行分类。在四个不同单细胞测序技术生成的数据集上测试,EnProCell性能优于现有先进方法。在参考数据集上预测准确率98.91%,F1分数98.64%;在未知细胞类型的查询数据上,准确率达99.52%,F1分数99.07%。此外,该方法结构简洁,无需额外计算资源,已开源于https://github.com/umar1196/EnProCell。
原文摘要 · Abstract (English)
Single-cell RNA sequencing (scRNA-seq) enables the study of cellular diversity at single cell level. It provides a global view of cell-type specification during the onset of biological mechanisms such as developmental processes and human organogenesis. Various statistical, machine and deep learning-based methods have been proposed for cell-type classification. Most of the methods utilizes unsupervised lower dimensional projections obtained from for a large reference data. In this work, we proposed a reference-based method for cell type classification, called EnProCell. The EnProCell, first, computes lower dimensional projections that capture both the high variance and class separability through an ensemble of principle component analysis and multiple discriminant analysis. In the second phase, EnProCell trains a deep neural network on the lower dimensional representation of data to classify cell types. The proposed method outperformed the existing state-of-the-art methods when tested on four different data sets produced from different single-cell sequencing technologies. The EnProCell showed higher accuracy (98.91) and F1 score (98.64) than other methods for predicting reference from reference datasets. Similarly, EnProCell also showed better performance than existing methods in predicting cell types for data with unknown cell types (query) from reference datasets (accuracy:99.52; F1 score: 99.07). In addition to improved performance, the proposed methodology is simple and does not require more computational resources and time. the EnProCell is available at https://github.com/umar1196/EnProCell.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。