用质量分布替代密度分布,避免聚类中对密集簇的偏见。
Mass Distribution versus Density Distribution in the Context of Clustering
- 以质量分布代替传统密度分布作为聚类基础
- 新算法MMC可发现任意形状、大小和密度的簇
- 适合需要公平识别稀疏与密集簇的研究者
本文研究聚类场景下数据的两种基本描述:密度分布与质量分布。自统计学提出以来,密度分布一直是数据分布的默认描述方式。然而我们证明,无论采用何种聚类算法,密度分布都存在根本性局限——高密度偏倚。现有基于密度的聚类算法虽通过不同机制缓解此偏倚并取得一定成效,但使用密度分布的根本缺陷仍阻碍发现任意形状、大小和密度的簇。为此,本文提出一种新算法,基于质量分布最大化所有簇的总质量,称为质量最大化聚类(MMC)。该算法亦可调整为最大化总密度,以对比密度与质量分布的内在差异。相比密度最大化聚类,MMC在优化过程中不偏向密集簇,具备更公平的聚类能力。
原文摘要 · Abstract (English)
This paper investigates two fundamental descriptors of data, i.e., density distribution versus mass distribution, in the context of clustering. Density distribution has been the de facto descriptor of data distribution since the introduction of statistics. We show that density distribution has its fundamental limitation -- high-density bias, irrespective of the algorithms used to perform clustering. Existing density-based clustering algorithms have employed different algorithmic means to counter the effect of the high-density bias with some success, but the fundamental limitation of using density distribution remains an obstacle to discovering clusters of arbitrary shapes, sizes and densities. Using the mass distribution as a better foundation, we propose a new algorithm which maximizes the total mass of all clusters, called mass-maximization clustering (MMC). The algorithm can be easily changed to maximize the total density of all clusters in order to examine the fundamental limitation of using density distribution versus mass distribution. The key advantage of the MMC over the density-maximization clustering is that the maximization is conducted without a bias towards dense clusters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。