改进哈廷根k均值算法,提升聚类效果2%-5%。
An effective variant of the Hartigan $k$-means algorithm
- 在哈廷根算法基础上微调,提升聚类质量。
- 实验显示性能提升2%-5%,维度或类别数越高越明显。
- 适合需要更高精度聚类的科研与工程应用。
k均值问题是经典的聚类问题,常与Lloyd算法(1957)等同。研究表明,哈廷根算法(1975)在几乎所有情况下表现更优,据Telgarsky-Vattani指出,典型性能提升为5%–10%。本文指出,对哈廷根方法进行一个极小的修改,可再带来2%–5%的性能提升;当数据维度或聚类数量k增加时,该优势更加显著。
原文摘要 · Abstract (English)
The k-means problem is perhaps the classical clustering problem and often synonymous with Lloyd's algorithm (1957). It has become clear that Hartigan's algorithm (1975) gives better results in almost all cases, Telgarsky-Vattani note a typical improvement of $5\%$ -- $10\%$. We point out that a very minor variation of Hartigan's method leads to another $2\%$ -- $5\%$ improvement; the improvement tends to become larger when either dimension or $k$ increase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。