arXiv:2502.00380cs.LGstat.ML2025-02

让现有聚类方法突破计算瓶颈,还能生成可解释的层次结构。

CoHiRF: Hierarchical Consensus for Interpretable Clustering Beyond Scalability Limits

  • 通过多视角共识与代表性收缩,分层提升聚类稳定性。
  • 在高维噪声和随机波动下仍保持稳定,支持原方法无法处理的大规模数据。
  • 适合需要可解释层次结构的大型聚类任务,如生物数据或图像分析。

我们提出 CoHiRF(Consensus Hierarchical Random Features),一种分层共识元算法,使现有聚类方法突破其固有的计算与内存限制。CoHiRF仅基于基础聚类方法输出的标签分配进行操作,不改变其目标函数、优化过程或几何假设。它反复在低维特征视图或随机实现上应用基础方法,通过共识机制强制一致,并利用基于代表的收缩逐步降低问题规模。在多种合成与真实数据集上,涵盖基于中心点、核函数、密度和图的聚类方法,结果表明 CoHiRF 可显著提升对高维噪声的鲁棒性,增强在随机变异下的稳定性,并使原方法在原本不可行的规模下仍可运行。我们还实证分析了分层共识的适用条件,强调可重现的标签关系及其与代表性收缩的兼容性。除扁平聚类外,CoHiRF 还生成显式的聚类融合层级,提供多分辨率且可解释的聚类结构。这些成果将分层共识定位为大规模聚类中实用且灵活的工具,无需修改原有方法行为即可扩展其适用范围。

原文摘要 · Abstract (English)

We introduce CoHiRF (Consensus Hierarchical Random Features), a hierarchical consensus framework that enables existing clustering methods to operate beyond their usual computational and memory limits. CoHiRF is a meta-algorithm that operates exclusively on the label assignments produced by a base clustering method, without modifying its objective function, optimization procedure, or geometric assumptions. It repeatedly applies the base method to multiple low-dimensional feature views or stochastic realizations, enforces agreement through consensus, and progressively reduces the problem size via representative-based contraction. Across a diverse set of synthetic and real-world experiments involving centroid-based, kernel-based, density-based, and graph-based methods, we show that CoHiRF can improve robustness to high-dimensional noise, enhance stability under stochastic variability, and enable scalability to regimes where the base method alone is infeasible. We also provide an empirical characterization of when hierarchical consensus is beneficial, highlighting the role of reproducible label relations and their compatibility with representative-based contraction. Beyond flat partitions, CoHiRF produces an explicit Cluster Fusion Hierarchy, offering a multi-resolution and interpretable view of the clustering structure. Together, these results position hierarchical consensus as a practical and flexible tool for large-scale clustering, extending the applicability of existing methods without altering their underlying behavior.

聚类层次结构可解释性大规模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。