在双曲空间中改进谱聚类,提升复杂层次数据的聚类效率。
Consistent Spectral Clustering in Hyperbolic Spaces
- 用双曲相似性矩阵替代欧氏空间的相似性矩阵。
- 实验表明双曲谱聚类收敛速度不低于欧氏版本。
- 适合处理树状、层级结构的数据,如生物分类或社交网络。
聚类作为无监督学习技术,在各类数据分析中扮演关键角色。尽管欧氏空间上的谱聚类已被广泛研究,但面对日益复杂的数据,欧氏空间在表示与学习上效率不足。尽管深度神经网络在双曲空间中受到关注,但非欧空间下的聚类算法或非深度机器学习模型仍研究不足。本文提出一种双曲空间上的谱聚类算法,以填补这一空白。双曲空间能更高效地表示层次化、树状结构数据。本方法将欧氏相似性矩阵替换为合适的双曲相似性矩阵,显著提升聚类效率。我们提出了双曲空间谱聚类算法,并证明其弱一致性,且收敛速度至少不慢于欧氏空间谱聚类。在威斯康星乳腺癌数据集上的实验验证了该方法优于传统欧氏方法。本工作为非欧空间在聚类中的应用开辟新路径,为处理复杂数据结构提供了新视角。
原文摘要 · Abstract (English)
Clustering, as an unsupervised technique, plays a pivotal role in various data analysis applications. Among clustering algorithms, Spectral Clustering on Euclidean Spaces has been extensively studied. However, with the rapid evolution of data complexity, Euclidean Space is proving to be inefficient for representing and learning algorithms. Although Deep Neural Networks on hyperbolic spaces have gained recent traction, clustering algorithms or non-deep machine learning models on non-Euclidean Spaces remain underexplored. In this paper, we propose a spectral clustering algorithm on Hyperbolic Spaces to address this gap. Hyperbolic Spaces offer advantages in representing complex data structures like hierarchical and tree-like structures, which cannot be embedded efficiently in Euclidean Spaces. Our proposed algorithm replaces the Euclidean Similarity Matrix with an appropriate Hyperbolic Similarity Matrix, demonstrating improved efficiency compared to clustering in Euclidean Spaces. Our contributions include the development of the spectral clustering algorithm on Hyperbolic Spaces and the proof of its weak consistency. We show that our algorithm converges at least as fast as Spectral Clustering on Euclidean Spaces. To illustrate the efficacy of our approach, we present experimental results on the Wisconsin Breast Cancer Dataset, highlighting the superior performance of Hyperbolic Spectral Clustering over its Euclidean counterpart. This work opens up avenues for utilizing non-Euclidean Spaces in clustering algorithms, offering new perspectives for handling complex data structures and improving clustering efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。