arXiv:2507.08311cs.LG2025-07

提出高效方法确定K均值聚类最优k值,速度提升99%。

CAS Condensed and Accelerated Silhouette: An Efficient Method for Determining the Optimal K in K-Means Clustering

  • 基于压缩轮廓系数与多种统计指标联合判断最优k
  • 高维数据下计算速度提升99%,精度与可扩展性兼顾
  • 适合实时聚类或资源受限场景的高效聚类需求

聚类是当前数据驱动环境中决策的关键组成部分,广泛应用于生物信息学、社交网络分析和图像处理等领域。然而,大规模数据集上的聚类精度仍是主要挑战。本文系统梳理了选择聚类最优k值的策略,重点在复杂数据环境下实现聚类精度与计算效率的平衡。针对文本和图像数据,提出改进的聚类技术以提升计算性能与聚类有效性。所提方法基于压缩轮廓系数(Condensed Silhouette),结合局部结构、间隙统计量、类一致性比和基于CCR与COI的重叠指数算法,自动计算K均值聚类的最优k值。对比实验表明,在高维数据集上,该方法执行速度最高提升99%,同时保持精度与可扩展性,适用于实时聚类或资源消耗最小化的应用场景。

原文摘要 · Abstract (English)

Clustering is a critical component of decision-making in todays data-driven environments. It has been widely used in a variety of fields such as bioinformatics, social network analysis, and image processing. However, clustering accuracy remains a major challenge in large datasets. This paper presents a comprehensive overview of strategies for selecting the optimal value of k in clustering, with a focus on achieving a balance between clustering precision and computational efficiency in complex data environments. In addition, this paper introduces improvements to clustering techniques for text and image data to provide insights into better computational performance and cluster validity. The proposed approach is based on the Condensed Silhouette method, along with statistical methods such as Local Structures, Gap Statistics, Class Consistency Ratio, and a Cluster Overlap Index CCR and COIbased algorithm to calculate the best value of k for K-Means clustering. The results of comparative experiments show that the proposed approach achieves up to 99 percent faster execution times on high-dimensional datasets while retaining both precision and scalability, making it highly suitable for real time clustering needs or scenarios demanding efficient clustering with minimal resource utilization.

聚类K均值高效算法参数选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。