用分布作为点的排斥型模型,提升聚类可解释性
Distributional Determinantal Point Process for Repulsive Clustering of Distributions

- 以分布为原子构建排斥点过程,基于切片Wasserstein核
- 在离散情形下证明了估计器的集中性,支持统计可靠性
- 适用于单细胞基因数据等复杂分布聚类,结果清晰可解释
我们提出分布型确定性点过程(dDPP),一种以概率分布为原子而非实数空间中点的新颖排斥型点过程。dDPP通过具有切片Wasserstein(SW)核的L-ensemble构造,并证明其为良定义的点过程。在离散设定下,我们推导出从分布原子中独立同分布采样时,L-ensemble、相关核及其行列式的插值估计量的集中性结果。基于此框架,我们提出一个分布值随机划分模型,采用排斥型广义贝叶斯混合模型:在混合测度的原子上施加dDPP先验,并基于SW距离定义广义似然。为总结后验推断,我们设计了一种决策论方法,以层级最优传输效用函数下的贝叶斯规则报告混合测度的点估计。该效用函数自然契合混合测度本身是分布之分布的性质。我们用该框架对单细胞基因表达数据和人类癫痫数据进行推断,得到反映数据真实结构的可解释且分离良好的聚类结果。
原文摘要 · Abstract (English)
We introduce the distributional determinantal point process (dDPP) as a novel repulsive point process whose atoms are probability distributions rather than points in a real space. The dDPP is constructed via an L-ensemble with a sliced Wasserstein (SW) kernel between distributions. We show its validity as a well-defined point process. In the discrete setting, we derive concentration results for plug-in estimators of the L-ensemble, the correlation kernel, and their determinants given i.i.d. samples from the distributional atoms. Leveraging this framework, we propose a distribution-valued random partition model by way of a repulsive generalized Bayesian mixture model. The model places a dDPP prior over the atoms of the mixing measure and defines a generalized likelihood based on SW distance. To summarize posterior inference, we develop a decision-theoretic approach to report a point estimate of the mixing measure as a Bayes rule under a hierarchical optimal transport utility function. The latter is a natural choice given that the mixing measure is itself a distribution over distributions. We use the proposed framework for inference with single-cell gene expression data and human epilepsy data, producing interpretable and well-separated clusters that reflect meaningful structure in the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。