arXiv:2509.18037stat.MLcs.LG2025-09被引 2

用核方法将概率分布映射到希尔伯特空间做聚类,适合高维数据。

Kernel K-means clustering of distributional data

  • 将每个概率分布映射为核均值嵌入,再在希尔伯特空间中进行K-means聚类。
  • 在合成孔径雷达图像上验证,能有效区分不同分布类型。
  • 方法简单高效,适用于高维(p>1)分布数据聚类,适合无监督分析。

我们研究从ℝ^p上的随机分布中采样的概率分布聚类问题。所提划分方法利用对称正定核k及其对应的再生核希尔伯特空间(RKHS)ℋ。通过将每个分布映射到ℋ中的核均值嵌入,得到ℋ中的样本,进而执行K-means聚类,实现对原始样本的无监督分类。该方法在p>1的高维情形下仍具计算可行性且实现简单。模拟研究揭示了核函数及其调参的选择影响。在一组合成孔径雷达(SAR)图像数据集上的实验展示了该聚类方法的有效性。

原文摘要 · Abstract (English)

We consider the problem of clustering a sample of probability distributions from a random distribution on $\mathbb R^p$. Our proposed partitioning method makes use of a symmetric, positive-definite kernel $k$ and its associated reproducing kernel Hilbert space (RKHS) $\mathcal H$. By mapping each distribution to its corresponding kernel mean embedding in $\mathcal H$, we obtain a sample in this RKHS where we carry out the $K$-means clustering procedure, which provides an unsupervised classification of the original sample. The procedure is simple and computationally feasible even for dimension $p>1$. The simulation studies provide insight into the choice of the kernel and its tuning parameter. The performance of the proposed clustering procedure is illustrated on a collection of Synthetic Aperture Radar (SAR) images.

聚类核方法分布数据SAR图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。