无需先验信息,通过中心匹配与边界过滤自动确定聚类数
CNMBI: Determining the Number of Clusters Using Center Pairwise Matching and Boundary Filtering
- 基于中心点位置动态比较,不依赖完整聚类结果
- 在CIFAR-10、STL-10等数据集上优于现有方法
- 首次引入样本置信度筛选,适合高维复杂数据
数据挖掘中的主要挑战之一是在无先验信息下确定最优聚类数。现有方法多基于聚类验证思想,通常对数据分布有隐含假设,难以适用于大规模图像和真实世界高维数据。为此,我们提出CNMBI方法。该方法利用数据空间内在分布信息,将聚类数确定转化为中心点间位置行为的动态比较过程,无需依赖完整聚类结果或设计复杂有效性指标。通过二分图理论高效建模此过程,并发现不同样本具有不同置信度,首次在聚类数判定中主动剔除低置信样本。CNMBI具有强鲁棒性,可灵活处理不同维度与形状的数据(如CIFAR-10、STL-10)。在多种挑战性数据集上的广泛对比实验表明,本方法显著优于当前最先进方法。
原文摘要 · Abstract (English)
One of the main challenges in data mining is choosing the optimal number of clusters without prior information. Notably, existing methods are usually in the philosophy of cluster validation and hence have underlying assumptions on data distribution, which prevents their application to complex data such as large-scale images and high-dimensional data from the real world. In this regard, we propose an approach named CNMBI. Leveraging the distribution information inherent in the data space, we map the target task as a dynamic comparison process between cluster centers regarding positional behavior, without relying on the complete clustering results and designing the complex validity index as before. Bipartite graph theory is then employed to efficiently model this process. Additionally, we find that different samples have different confidence levels and thereby actively remove low-confidence ones, which is, for the first time to our knowledge, considered in cluster number determination. CNMBI is robust and allows for more flexibility in the dimension and shape of the target data (e.g., CIFAR-10 and STL-10). Extensive comparison studies with state-of-the-art competitors on various challenging datasets demonstrate the superiority of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。