通过最大化标签的Schatten p-范数,实现数据结构与类别标签的一致性聚类。
Manifold Clustering with Schatten p-norm Maximization
- 用标签引导流形结构,确保聚类结果与数据内在结构一致。
- 通过最大化Schatten p-范数,在聚类中自然保持类别平衡。
- 兼容多种距离函数,适用于非线性可分数据,灵活性强。
流形聚类因其强大的复杂数据结构捕捉能力,在聚类分析中占据重要地位。然而,现有方法通常只关注K-means与流形学习之间的最优组合,忽视了数据结构与标签之间的一致性。为此,我们深入探讨了K-means与流形学习的关系,基于此融合二者,提出一种新的聚类框架。算法利用标签引导流形结构,并在该结构上进行聚类,确保数据结构与标签的一致性。此外,为在聚类过程中自然维持类别平衡,我们最大化标签的Schatten p-范数,并提供了理论证明。同时,该聚类框架设计灵活,兼容多种距离函数,可高效处理非线性可分数据。多个数据库的实验结果验证了所提模型的优越性。
原文摘要 · Abstract (English)
Manifold clustering, with its exceptional ability to capture complex data structures, holds a pivotal position in cluster analysis. However, existing methods often focus only on finding the optimal combination between K-means and manifold learning, and overlooking the consistency between the data structure and labels. To address this issue, we deeply explore the relationship between K-means and manifold learning, and on this basis, fuse them to develop a new clustering framework. Specifically, the algorithm uses labels to guide the manifold structure and perform clustering on it, which ensures the consistency between the data structure and labels. Furthermore, in order to naturally maintain the class balance in the clustering process, we maximize the Schatten p-norm of labels, and provide a theoretical proof to support this. Additionally, our clustering framework is designed to be flexible and compatible with many types of distance functions, which facilitates efficient processing of nonlinear separable data. The experimental results of several databases confirm the superiority of our proposed model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。