用图传播重构密度聚类,让高维数据聚类更快更准。
Towards Robust and Scalable Density-based Clustering via Graph Propagation

- 将聚类转为图上的标签传播,自动处理不同密度区域。
- 可处理百万级数据点,分钟内完成,精度超越现有方法。
- 不依赖距离度量,适合大规模真实数据场景。
我们提出一种名为 CluProp 的新框架,将高维空间中的多密度聚类重新构想为邻域图上的标签传播过程。该方法在密度聚类与图连通性之间建立正式关联,利用网络科学中的高效传播机制,缓解传统密度聚类对参数的敏感性。具体而言,引入确定性的密度驱动传播策略,实现可扩展的邻域识别。该框架对距离度量选择无感,在大规模数据上表现优异,可在几分钟内处理数百万个点,并在准确率上持续优于现有基线方法。
原文摘要 · Abstract (English)
We present \textit{CluProp}, a novel framework that reimagines varied-density clustering in high-dimensional spaces as a label propagation process over neighborhood graphs. Our approach formally bridges the gap between density-based clustering and graph connectivity, leveraging efficient propagation mechanisms from network science to mitigate the parameter sensitivity inherent in traditional density-based methods. Specifically, we introduce a deterministic density-based propagation strategy to ensure scalable neighborhood identification. The framework is agnostic to the choice of distance metric and exhibits superior performance on large-scale data, processing millions of points in minutes while consistently outperforming existing baselines in accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。