用核方法在希尔伯特空间建模高斯混合,解决无限维数据聚类难题
Gaussian mixture models in Hilbert spaces via kernel methods

- 基于核均值嵌入构建希尔伯特空间上的高斯混合模型
- 理论证明算法有效,可逼近无限维空间中的任意概率分布
- 适用于动态函数数据、图结构等复杂医学数据的聚类分析
许多学科的现代数据包含随时间演变、可能为无限维的随机对象,如动态函数数据,自然应建模在希尔伯特空间中。在此类场景下,通过密度表征概率测度往往不成立或技术上困难。针对聚类应用,我们提出基于核均值嵌入的希尔伯特空间值数据高斯混合框架,并开发高效的估计优化算法。理论证明该算法定义良好,且模型在无限维空间中可生成稠密近似类。我们在多种结构与数据几何上进行大量实验评估,包括 $L^2$-函数数据和现代医疗应用中的拉普拉斯空间随机图。
原文摘要 · Abstract (English)
Modern datasets across many disciplines increasingly consist of time-evolving, potentially infinite-dimensional random objects, such as dynamic functional data, which are naturally modeled in Hilbert spaces. In these settings, characterizing probability measures, for example, through densities, can be ill-defined or technically challenging. Motivated by clustering applications, we propose a Gaussian mixture framework for Hilbert-space-valued data based on kernel mean embeddings and develop efficient optimization algorithms for estimation. We establish theoretical guarantees showing that the proposed algorithm is well defined and that the model yields a dense class of approximations in infinite-dimensional spaces. We evaluate the framework through extensive experiments on diverse structures and data geometries, including $L^2$-functional data and random graphs in Laplacian spaces arising in modern medical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。