用中性逻辑提升聚类精度与可解释性,适合处理含不确定性的复杂数据。
UNCA: A Neutrosophic-Based Framework for Robust Clustering and Enhanced Data Interpretation
- 基于中性逻辑构建聚类框架,融合真值、不确定性与假值三重度量。
- 在多个数据集上表现优于传统方法,如鸢尾花数据集轮廓系数达0.89。
- 支持动态网络可视化,增强聚类结果的直观解释,适合高维数据探索。
准确刻画海量数据中复杂的关联关系与内在不确定性仍是数据聚类领域的重大挑战。本文提出统一中性聚类算法(UNCA),结合多维度策略与中性逻辑以提升聚类性能。UNCA首先通过λ-切割矩阵进行全量相似性分析,筛选出有意义的数据点间关系;随后基于真值、不确定性与假值的隶属度初始化中性K均值聚类中心。算法融合动态网络可视化与最小生成树(MST),清晰呈现聚类间的关联结构。采用单值中性集(SVNS)优化聚类分配,经模糊化相似度度量后确保精确聚类结果,并通过去模糊化方法固化最终聚类标签。实验表明,UNCA在多个指标上超越传统方法:在鸢尾花数据集上轮廓系数达0.89,在葡萄酒数据集上戴维斯-博尔丁指数为0.59,在数字数据集上调整兰德指数(ARI)为0.76,在客户细分数据集上标准化互信息(NMI)为0.80。结果证明,相比模糊C均值(FCM)、中性C均值(NCM)及核中性C均值(KNCM),UNCA显著提升聚类准确性、可解释性与鲁棒性,适用于复杂数据处理任务。
原文摘要 · Abstract (English)
Accurately representing the complex linkages and inherent uncertainties included in huge datasets is still a major difficulty in the field of data clustering. We address these issues with our proposed Unified Neutrosophic Clustering Algorithm (UNCA), which combines a multifaceted strategy with Neutrosophic logic to improve clustering performance. UNCA starts with a full-fledged similarity examination via a λ-cutting matrix that filters meaningful relationships between each two points of data. Then, we initialize centroids for Neutrosophic K-Means clustering, where the membership values are based on their degrees of truth, indeterminacy and falsity. The algorithm then integrates with a dynamic network visualization and MST (Minimum Spanning Tree) so that a visual interpretation of the relationships between the clusters can be clearly represented. UNCA employs SingleValued Neutrosophic Sets (SVNSs) to refine cluster assignments, and after fuzzifying similarity measures, guarantees a precise clustering result. The final step involves solidifying the clustering results through defuzzification methods, offering definitive cluster assignments. According to the performance evaluation results, UNCA outperforms conventional approaches in several metrics: it achieved a Silhouette Score of 0.89 on the Iris Dataset, a Davies-Bouldin Index of 0.59 on the Wine Dataset, an Adjusted Rand Index (ARI) of 0.76 on the Digits Dataset, and a Normalized Mutual Information (NMI) of 0.80 on the Customer Segmentation Dataset. These results demonstrate how UNCA enhances interpretability and resilience in addition to improving clustering accuracy when contrasted with Fuzzy C-Means (FCM), Neutrosophic C-Means (NCM), as well as Kernel Neutrosophic C-Means (KNCM). This makes UNCA a useful tool for complex data processing tasks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。