arXiv:2411.00268cs.LG2024-11

通过高阶一致性学习融合多源聚类结果,提升聚类精度。

Clustering ensemble algorithm with high-order consistency learning

  • 基于数据内在关联构建高阶信息融合矩阵
  • 相比最优基线算法,准确率提升7.22%,NMI提升9.19%
  • 适合处理低质量基聚类影响的复杂数据聚类任务

现有聚类集成研究多聚焦于设计实用的一致性学习算法。为解决基聚类质量参差不齐、劣质基聚类影响集成性能的问题,本文从数据挖掘角度,基于基聚类挖掘数据内在联系,提出一种高阶信息融合算法——高阶一致性学习聚类集成(HCLCE)。首先将各高阶信息融合为新的结构化一致性矩阵,再将多个一致性矩阵进行融合,最终实现多信息一致性输出。实验表明,与次优的局部加权证据累积(LWEA)算法相比,所提方法在聚类准确率上平均提升7.22%,标准化互信息(NMI)平均提升9.19%。结果表明,该算法在聚类集成效果上优于单一信息或传统集成方法。

原文摘要 · Abstract (English)

Most of the research on clustering ensemble focuses on designing practical consistency learning algorithms.To solve the problems that the quality of base clusters varies and the low-quality base clusters have an impact on the performance of the clustering ensemble, from the perspective of data mining, the intrinsic connections of data were mined based on the base clusters, and a high-order information fusion algorithm was proposed to represent the connections between data from different dimensions, namely Clustering Ensemble with High-order Consensus learning (HCLCE). Firstly, each high-order information was fused into a new structured consistency matrix. Then, the obtained multiple consistency matrices were fused together. Finally, multiple information was fused into a consistent result. Experimental results show that LCLCE algorithm has the clustering accuracy improved by an average of 7.22%, and the Normalized Mutual Information (NMI) improved by an average of 9.19% compared with the suboptimal Locally Weighted Evidence Accumulation (LWEA) algorithm. It can be seen that the proposed algorithm can obtain better clustering results compared with clustering ensemble algorithms and using one information alone.

聚类集成高阶一致性信息融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。