arXiv:2503.14357cs.LGstat.AP2025-03被引 2

用瓦瑟斯坦距离结合核方法,实现对复杂数据的高效聚类。

Wasserstein-based Kernel Principal Component Analysis for Clustering Applications

  • 基于多个参考分布近似计算瓦瑟斯坦距离,提升效率。
  • 通过平移正定核与核主成分分析,映射分布数据到特征空间。
  • 适合处理向量和分布数据,适用于电力图谱、时间序列等场景。

许多数据聚类应用需处理无法表示为向量的复杂对象。在此背景下,词袋向量表示通过离散分布描述复杂对象,瓦瑟斯坦距离提供了良好的相似性度量。核方法通过将距离信息嵌入特征空间来促进分析。然而,目前仍缺乏将核方法与瓦瑟斯坦距离结合用于分布数据聚类的无监督框架。本文提出一个计算可处理的框架,整合瓦瑟斯坦度量与核方法进行聚类。该框架包含三部分:(i) 使用多个参考分布高效近似成对瓦瑟斯坦距离;(ii) 基于瓦瑟斯坦距离的平移正定核函数,结合核主成分分析实现特征映射;(iii) 可扩展且距离无关的聚类有效性指标,用于评估与核参数优化。在电力分布图谱和真实时间序列数据上的实验表明,该框架在效果与效率上均表现优异。

原文摘要 · Abstract (English)

Many data clustering applications must handle objects that cannot be represented as vectors. In this context, the bag-of-vectors representation describes complex objects through discrete distributions, for which the Wasserstein distance provides a well-conditioned dissimilarity measure. Kernel methods extend this by embedding distance information into feature spaces that facilitate analysis. However, an unsupervised framework that combines kernels with Wasserstein distances for clustering distributional data is still lacking. We address this gap by introducing a computationally tractable framework that integrates Wasserstein metrics with kernel methods for clustering. The framework can accommodate both vectorial and distributional data, enabling applications in various domains. It comprises three components: (i) an efficient approximation of pairwise Wasserstein distances using multiple reference distributions; (ii) shifted positive definite kernel functions based on Wasserstein distances, combined with kernel principal component analysis for feature mapping; and (iii) scalable, distance-agnostic validity indices for clustering evaluation and kernel parameter optimization. Experiments on power distribution graphs and real-world time series demonstrate the effectiveness and efficiency of the proposed framework.

聚类瓦瑟斯坦核方法分布数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。