用分布视角改进分层聚类,让结果更符合真实结构。
Rethinking Divisive Hierarchical Clustering from a Distributional Perspective
- 改用分布核替代集合分割标准,避免随意分裂。
- 新方法在生物数据上实现更高总相似度,达理论下界。
- 适合生物信息学等需真实结构对齐的场景。
现有基于目标的分裂式分层聚类(DHC)方法生成的树状图缺乏三个理想特性:无不当分裂、相似聚类应归入同一子集、与真实标签对应。根源在于使用集合导向的二分评估准则。本文提出采用分布核替代该准则,实现以最大化所有聚类总相似度(TSC)为目标的新分布导向方法。理论分析表明,所得树状图保证了TSC的下界。实验验证了该方法在人工数据和空间转录组学(bioinformatics)数据上的有效性。在空间转录组学数据中,本方法成功构建出与生物学区域一致的树状图,而其他方法均失败。
原文摘要 · Abstract (English)
We uncover that current objective-based Divisive Hierarchical Clustering (DHC) methods produce a dendrogram that does not have three desired properties i.e., no unwarranted splitting, group similar clusters into a same subset, ground-truth correspondence. This shortcoming has their root cause in using a set-oriented bisecting assessment criterion. We show that this shortcoming can be addressed by using a distributional kernel, instead of the set-oriented criterion; and the resultant clusters achieve a new distribution-oriented objective to maximize the total similarity of all clusters (TSC). Our theoretical analysis shows that the resultant dendrogram guarantees a lower bound of TSC. The empirical evaluation shows the effectiveness of our proposed method on artificial and Spatial Transcriptomics (bioinformatics) datasets. Our proposed method successfully creates a dendrogram that is consistent with the biological regions in a Spatial Transcriptomics dataset, whereas other contenders fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。