arXiv:2504.03928cs.LGmath.PR2025-04被引 7

用概率距离函数替代传统距离,让聚类适应随机数据。

Random Normed k-Means: A Paradigm-Shift in Clustering within Probabilistic Metric Spaces

  • 以距离分布函数代替固定距离,构建概率度量空间下的新聚类范式。
  • 在多种真实与合成数据上表现优于k-means++等经典方法。
  • 适合处理非线性可分、随机性强的数据,为未来研究提供理论基础。

现有方法多受限于传统距离度量,难以有效处理随机数据。本文首次提出一种在概率度量空间中运行的k-means变体,将传统距离度量替换为定义良好的距离分布函数。该方法在确定性和随机数据集上均表现出更强的灵活性与鲁棒性,为随机环境中的聚类研究奠定了新基础。通过采用概率视角,不仅引入全新范式,还建立了严谨的理论框架,有望成为未来涉及随机数据聚类研究的关键参考。在多样真实与合成数据集上的实验表明,所提随机归一化k-means(RNKM)算法在轮廓系数、Davies-Bouldin、Calinski-Harabasz、调整兰德指数和重构误差等主流评估指标上均优于k-means++、模糊c均值及核概率k-means等方法。特别地,RNKM展现出识别非线性可分结构的强大能力,适用于复杂聚类场景。该成果标志着聚类研究的重大突破,为传统技术提供了有力替代方案,并填补了文献中的长期空白。通过连接概率度量与聚类,本研究为动态数据驱动应用中的高级数据分析开辟了新路径。

原文摘要 · Abstract (English)

Existing approaches remain largely constrained by traditional distance metrics, limiting their effectiveness in handling random data. In this work, we introduce the first k-means variant in the literature that operates within a probabilistic metric space, replacing conventional distance measures with a well-defined distance distribution function. This pioneering approach enables more flexible and robust clustering in both deterministic and random datasets, establishing a new foundation for clustering in stochastic environments. By adopting a probabilistic perspective, our method not only introduces a fresh paradigm but also establishes a rigorous theoretical framework that is expected to serve as a key reference for future clustering research involving random data. Extensive experiments on diverse real and synthetic datasets assess our model's effectiveness using widely recognized evaluation metrics, including Silhouette, Davies-Bouldin, Calinski Harabasz, the adjusted Rand index, and distortion. Comparative analyses against established methods such as k-means++, fuzzy c-means, and kernel probabilistic k-means demonstrate the superior performance of our proposed random normed k-means (RNKM) algorithm. Notably, RNKM exhibits a remarkable ability to identify nonlinearly separable structures, making it highly effective in complex clustering scenarios. These findings position RNKM as a groundbreaking advancement in clustering research, offering a powerful alternative to traditional techniques while addressing a long-standing gap in the literature. By bridging probabilistic metrics with clustering, this study provides a foundational reference for future developments and opens new avenues for advanced data analysis in dynamic, data-driven applications.

聚类概率度量随机数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。