arXiv:2512.11448cs.LGstat.ML2025-12

将均值漂移算法拓展到双曲空间,更好发现树状数据结构。

Hyperbolic Gaussian Blurring Mean Shift: A Statistical Mode-Seeking Framework for Clustering in Curved Spaces

  • 用双曲距离和莫比乌斯加权均值,保持几何一致性。
  • 在11个真实数据集上显著提升树状结构聚类效果。
  • 适合处理具有层次关系的复杂数据,如生物分类、社交网络。

聚类是揭示数据模式的基本无监督学习任务。尽管高斯模糊均值漂移(GBMS)在欧几里得空间中能有效识别任意形状的聚类,但在具有层次或树状结构的数据上表现不佳。本文提出HypeGBMS,一种将GBMS扩展至双曲空间的新方法。该方法用双曲距离替代欧氏计算,并采用莫比乌斯加权均值确保所有更新与空间几何一致。HypeGBMS能有效捕捉潜在层次结构,同时保留GBMS的密度搜索特性。我们提供了收敛性和计算复杂度的理论分析,并通过实证结果验证了其在层次数据上的聚类质量提升。该工作连接经典均值漂移与双曲表示学习,为弯曲空间中的密度聚类提供了一个严谨的方法。在11个真实世界数据集上的广泛实验表明,HypeGBMS在非欧几里得场景下显著优于传统均值漂移方法,凸显其鲁棒性与有效性。

原文摘要 · Abstract (English)

Clustering is a fundamental unsupervised learning task for uncovering patterns in data. While Gaussian Blurring Mean Shift (GBMS) has proven effective for identifying arbitrarily shaped clusters in Euclidean space, it struggles with datasets exhibiting hierarchical or tree-like structures. In this work, we introduce HypeGBMS, a novel extension of GBMS to hyperbolic space. Our method replaces Euclidean computations with hyperbolic distances and employs Möbius-weighted means to ensure that all updates remain consistent with the geometry of the space. HypeGBMS effectively captures latent hierarchies while retaining the density-seeking behavior of GBMS. We provide theoretical insights into convergence and computational complexity, along with empirical results that demonstrate improved clustering quality in hierarchical datasets. This work bridges classical mean-shift clustering and hyperbolic representation learning, offering a principled approach to density-based clustering in curved spaces. Extensive experimental evaluations on $11$ real-world datasets demonstrate that HypeGBMS significantly outperforms conventional mean-shift clustering methods in non-Euclidean settings, underscoring its robustness and effectiveness.

聚类双曲空间密度聚类层次结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。