用KNN图动态评估标签数据安全性,提升半监督聚类准确率
K-GBS3FCM -- KNN Graph-Based Safe Semi-Supervised Fuzzy C-Means
- 基于KNN图分析标签与无标签数据的邻域关系
- 在56组实验中64%情况下优于其他方法,提升聚类精度
- 适合需利用少量标签数据增强聚类效果的研究者
利用先验领域知识(部分标签数据)进行聚类近年来受到广泛关注,称为半监督聚类。为确保先验知识的安全性,提出安全半监督聚类(S3C)算法。本文提出基于K近邻(KNN)图的安全感知半监督模糊c均值算法(K-GBS3FCM),通过KNN动态评估标签与无标签数据间的邻域关系,优化标签数据使用并降低错误标签影响。引入正则化参数与平均安全度机制调节标签对无标签数据的影响。在多个基准数据集上的实验表明,该图方法有效提升聚类准确性,在56组测试配置中,64%的情况下性能显著优于其他半监督及传统无监督方法。研究展示了将KNN等图方法与经典聚类技术结合的潜力,适用于需融合标签与无标签数据的场景。
原文摘要 · Abstract (English)
Clustering data using prior domain knowledge, starting from a partially labeled set, has recently been widely investigated. Often referred to as semi-supervised clustering, this approach leverages labeled data to enhance clustering accuracy. To maximize algorithm performance, it is crucial to ensure the safety of this prior knowledge. Methods addressing this concern are termed safe semi-supervised clustering (S3C) algorithms. This paper introduces the KNN graph-based safety-aware semi-supervised fuzzy c-means algorithm (K-GBS3FCM), which dynamically assesses neighborhood relationships between labeled and unlabeled data using the K-Nearest Neighbors (KNN) algorithm. This approach aims to optimize the use of labeled data while minimizing the adverse effects of incorrect labels. Additionally, it is proposed a mechanism that adjusts the influence of labeled data on unlabeled ones through regularization parameters and the average safety degree. Experimental results on multiple benchmark datasets demonstrate that the graph-based approach effectively leverages prior knowledge to enhance clustering accuracy. The proposed method was significantly superior in 64% of the 56 test configurations, obtaining higher levels of clustering accuracy when compared to other semi-supervised and traditional unsupervised methods. This research highlights the potential of integrating graph-based approaches, such as KNN, with established techniques to develop advanced clustering algorithms, offering significant applications in fields that rely on both labeled and unlabeled data for more effective clustering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。