arXiv:2507.13417cs.LGcs.AI2025-07

提出Soft-ECM算法,让模糊聚类支持复杂数据

Soft-ECM: An extension of Evidential C-Means for complex data

  • 用半度量重构证据C均值,支持非欧空间数据
  • 在数值数据上效果接近传统模糊聚类,处理混合数据能力更强
  • 适合处理时间序列等复杂数据,尤其配合DTW度量时表现优异

基于信念函数的聚类方法因能有效表示不确定性和不精确性而受到机器学习领域关注。然而,现有算法无法应用于复杂数据(如数值与类别混合数据或时间序列等非表格数据)。这类数据通常不在欧氏空间中,而现有方法依赖欧氏空间特性来构建质心。本文重新构建了证据C均值(ECM)问题以适用于复杂数据,提出新算法Soft-ECM,仅需半度量即可一致地定位模糊聚类中心。实验表明,Soft-ECM在数值数据上的结果可媲美传统模糊聚类方法,并展示了其处理混合数据的能力,以及在结合模糊聚类与半度量(如DTW)处理时间序列数据时的优势。

原文摘要 · Abstract (English)

Clustering based on belief functions has been gaining increasing attention in the machine learning community due to its ability to effectively represent uncertainty and/or imprecision. However, none of the existing algorithms can be applied to complex data, such as mixed data (numerical and categorical) or non-tabular data like time series. Indeed, these types of data are, in general, not represented in a Euclidean space and the aforementioned algorithms make use of the properties of such spaces, in particular for the construction of barycenters. In this paper, we reformulate the Evidential C-Means (ECM) problem for clustering complex data. We propose a new algorithm, Soft-ECM, which consistently positions the centroids of imprecise clusters requiring only a semi-metric. Our experiments show that Soft-ECM present results comparable to conventional fuzzy clustering approaches on numerical data, and we demonstrate its ability to handle mixed data and its benefits when combining fuzzy clustering with semi-metrics such as DTW for time series data.

聚类模糊聚类复杂数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。