针对单细胞数据的层次化新类发现,提出改进的聚类方法。
Hierarchical novel class discovery for single-cell transcriptomic profiles
- 基于层次结构设计改进的k-Means与GMM聚类算法
- 在人工与真实数据上实现高精度聚类与标签映射
- 适用于发育生物学中带标签与无标签数据分离的场景
单细胞转录组学实验面临的核心挑战是如何标注单细胞转录组谱。由于数据规模大、维度高,亟需自动化标注方法。本文聚焦发育生物学背景下的数据,其分化过程具有天然的层次结构。考虑一种常见情形:训练时既有带标签数据,也有无标签数据,但两者的标签集互不重叠,属于新类发现问题。目标是同时完成数据聚类和将聚类结果映射到标签。本文提出k-Means与GMM聚类方法的扩展版本,并在人工生成及真实转录组数据集上报告了对比实验结果。所提方法充分利用了数据的层次特性。
原文摘要 · Abstract (English)
One of the major challenges arising from single-cell transcriptomics experiments is the question of how to annotate the associated single-cell transcriptomic profiles. Because of the large size and the high dimensionality of the data, automated methods for annotation are needed. We focus here on datasets obtained in the context of developmental biology, where the differentiation process leads to a hierarchical structure. We consider a frequent setting where both labeled and unlabeled data are available at training time, but the sets of the labels of labeled data on one side and of the unlabeled data on the other side, are disjoint. It is an instance of the Novel Class Discovery problem. The goal is to achieve two objectives, clustering the data and mapping the clusters with labels. We propose extensions of k-Means and GMM clustering methods for solving the problem and report comparative results on artificial and experimental transcriptomic datasets. Our approaches take advantage of the hierarchical nature of the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。