arXiv:2604.04302stat.MEcs.LG2026-04

提出无需调参的K-means合并算法,提升非线性数据聚类效果。

CavMerge: Merging K-means Based on Local Log-Concavity

  • 基于局部对数凹性设计聚类合并策略,避免传统方法依赖超参数
  • 在模拟与真实数据上均优于现有最优算法,结果更稳定可靠
  • 计算高效且理论保证强,适合追求鲁棒聚类的科研与工程应用

K-means聚类作为经典广泛使用的聚类方法,在处理非线性可分数据时表现不佳。为解决此问题,已有诸多改进方法被提出,包括通过较大K值的K-means结果合并获得最终聚类分配。然而,现有方法常面临计算效率低和超参数调优困难的问题。本文提出新的聚类合并算法CavMerge,具有直观性、无需参数调优且计算高效的特点。该算法在较弱的局部分布假设下仍具备强一致性与快速收敛性。在多种模拟与真实数据集上的实证研究表明,本方法生成的聚类结果比当前最先进算法更为可靠。

原文摘要 · Abstract (English)

K-means clustering, a classic and widely-used clustering technique, is known to exhibit suboptimal performance when applied to non-linearly separable data. Numerous adjustments and modifications have been proposed to address this issue, including methods that merge K-means results from a relatively large K to obtain a final cluster assignment. However, existing methods of this nature often encounter computational inefficiencies and suffer from hyperparameter tuning. Here we present \emph{CavMerge}, a novel K-means merging algorithm that is intuitive, free of parameter tuning, and computationally efficient. Operating under minimal local distributional assumptions, our algorithm demonstrates strong consistency and rapid convergence guarantees. Empirical studies on various simulated and real datasets demonstrate that our method yields more reliable clusters in comparison to current state-of-the-art algorithms.

聚类无监督学习算法优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。